AI Agents Full Course 2026: Master Agentic AI (2 Hours) — Transcript
Full transcript
- 0:00Hey, this is the definitive course on AI
- 0:02agents. I currently teach over 2,000
- 0:04people how to use AI agents in both
- 0:05their personal and business lives and
- 0:07run a business that does over $4 million
- 0:09a year using AI agents. So, you don't
- 0:11need any programming or pre-existing
- 0:13computer experience in order to make
- 0:14this course work for you. I myself don't
- 0:16have a formal computer science degree.
- 0:18I've learned everything that I know
- 0:19watching free resources like you doing
- 0:21now. This is also a general AI agents
- 0:23course, so you don't need to know any
- 0:25specific platform. This isn't just on
- 0:27Codex or Claude Code or Anti-Gravity,
- 0:29but rather on all of them. So, wherever
- 0:31you guys are starting, you'll end up at
- 0:33the same place. No fluff, here's what
- 0:34you're going to learn in this course.
- 0:35First, I'll show you guys a demo where
- 0:37I'm controlling five AI agents, each
- 0:39with their own Chrome browsers as they
- 0:41interact with the web and perform
- 0:42economically valuable activities for me.
- 0:44I wanted to front-load this course with
- 0:46a demo so you guys could see what we're
- 0:47working up to. And just a few months
- 0:48ago, what I'm doing here would have been
- 0:50considered absurd. Then, I'm going to
- 0:51cover the core AI agent workflow loop,
- 0:53which works independent of which
- 0:54platform you're using. After that, I'm
- 0:56actually going to talk about and then
- 0:57sign up to the three major AI agent
- 0:59platforms right now. So, I'll sign up to
- 1:00Codex, to Anti-Gravity, and then Claude
- 1:02Code. And then after I'll cover what
- 1:04each platform is at the moment the best
- 1:06or the worst at. Then, we're going to
- 1:07dive into foundational AI agent
- 1:09prompting techniques. So, self-modifying
- 1:11agent instructions where the agent will
- 1:13rewrite its own rules to minimize the
- 1:14number of errors made. Multi-agent MCP
- 1:17orchestration, which is where we'll
- 1:18register Codex, Gemini, and Claude as
- 1:20MCP servers so you can manage multiple
- 1:22agents within a single conversation
- 1:24thread. Video-to-action pipelines where
- 1:26we'll teach agents to learn from YouTube
- 1:28videos instead of plain text alone.
- 1:30Stochastic multi-agent consensus where
- 1:31we'll spawn agents with the same prompt
- 1:34and then use their statistical spread in
- 1:36order to ideate and improve things
- 1:37about. Agent chat rooms where you'll
- 1:39build centralized places for agents to
- 1:41debate ideas, pushing them to much
- 1:43higher quality answers than before.
- 1:45Sub-agent verification loops where your
- 1:47agents will actually review each other's
- 1:48work in real time to catch things that
- 1:50one of them might have missed.
- 1:52We'll talk prompt contracts. I'll show
- 1:53you guys reverse prompting and a bunch
- 1:55of other techniques as well. And
- 1:57finally, we'll chat about context
- 1:58management and improving the agent
- 2:00output quality before closing out by
- 2:02discussing how to optimize AI agent and
- 2:05then token pricing. So far, I haven't
- 2:06seen anybody on YouTube discuss most of
- 2:08what I cover in this course. So for all
- 2:09intents and purposes, you guys consider
- 2:11this the sauce. Please bookmark this
- 2:13video, subscribe to the channel, and
- 2:14let's get into it. First, I want to show
- 2:16you how powerful these agents can be
- 2:18when you learn how to distribute work
- 2:20across multiple Chrome instances and
- 2:22give each sub-agent their own workspace.
- 2:24What I have here is a simple list of
- 2:27leads from, let's just say, a
- 2:29conference. Now, we have fields like
- 2:31their websites, their LinkedIn
- 2:33description, their first name, their
- 2:34last name, but one thing is missing:
- 2:37their email address. Now, just a year
- 2:39ago or so, that would have invalidated
- 2:42my ability to reach out to these leads.
- 2:44But now, because I possess their
- 2:46websites, I can actually spawn a bunch
- 2:48of Claude code agents, have them go to
- 2:50the websites, then have them
- 2:51interactively and dynamically fill out
- 2:53their contact forms.
- 2:55So what just happened as I was talking
- 2:57was Claude went ahead and then opened up
- 2:59a bunch of different Chrome browsers for
- 3:00me.
- 3:01I'm going to rearrange these to make it
- 3:02really easy to see. And so, this might
- 3:04be a little bit tough to see, but what
- 3:06these agents are all doing is they're
- 3:07independently navigating over to the
- 3:10contact fields of each of these
- 3:11websites. They're then dynamically
- 3:13filling out fields like the first name,
- 3:16the last name, the email address, and so
- 3:18on and so forth. And then they're
- 3:19putting in a little bit of outreach
- 3:22that's templated, but then changes
- 3:23depending on who they're reaching out
- 3:25to. These agents, through a combination
- 3:26of both research and then communication
- 3:28between each other in a shared chat
- 3:30room, are capable of doing things that
- 3:32any one agent might have taken many,
- 3:34many hours to do before. This is what
- 3:35I'm going to work up to with you guys
- 3:37over the course of the rest of the next
- 3:39couple of hours. The main strength of AI
- 3:41agents is really their ability to
- 3:43parallelize, which is to run multiple
- 3:46instances of each of them simultaneously
- 3:48while they accomplish a task. Now, right
- 3:51now, I would say most AI agents aren't
- 3:53as intelligent or as capable as a human
- 3:55being for any given need. But, what they
- 3:58are much better at us than is being
- 4:00fast. And so, despite the fact that
- 4:03their accuracy might be a little bit
- 4:04lower than a human, their ability to
- 4:06one-shot stuff is worse than ours at the
- 4:09moment, they can run multiple instances
- 4:11of themselves simultaneously and try
- 4:13multiple approaches over and over and
- 4:15over and over again in order to
- 4:17ultimately achieve much better results
- 4:19than we can. The key is you need to know
- 4:21a little bit about how they work under
- 4:23the hood. Then, you need to be able to
- 4:24combine them using elaborate prompt
- 4:26architecture like I'm going to show you
- 4:28in this course. So, why don't we start
- 4:29with one of the simplest, most
- 4:30foundational concepts before I actually
- 4:32guide you guys through signing up and
- 4:34setting up these different agents. And I
- 4:36call this the core agent loop.
- 4:39To make a long story short, I think most
- 4:41of you probably have intuition about how
- 4:43agents do things, but really what
- 4:45they're doing at the end of the day is
- 4:48they're going through a loop over and
- 4:50over and over again. And this loop is
- 4:52composed of three major functions.
- 4:55The first is the observation step. And
- 4:58so, here the agent is basically reading
- 5:00through all of its context. We're going
- 5:02to chat a little bit more about how to
- 5:04optimize and manage that later. That
- 5:06includes things like its files, its
- 5:08previous tool calls, it includes all of
- 5:11the system prompts, the Claude, Gemini,
- 5:13and agents.mds that you provide. If it
- 5:16does research in a previous step, it'll
- 5:18include the research from the internet.
- 5:20Uh if you're feeding in multimodal data
- 5:22like vision data, camera data,
- 5:24uh you know, audio files, and so on and
- 5:26so forth, it'll include all of that. And
- 5:28so, this agent, okay, is just in an
- 5:30environment and it's just always
- 5:32observing what's going around it, at
- 5:34least to start, in the observation step.
- 5:37From there, it'll reason. And so, this
- 5:39is the think step. Here, it'll consider,
- 5:42based off of all of this context and
- 5:44based off of, you know, the user's
- 5:45high-level goal, what do I do next? How
- 5:47should I plan my approach? And nowadays,
- 5:50most agentic coding platforms make use
- 5:52of like a dedicated reasoning step that
- 5:55you can actually click into and see,
- 5:57which I'll show you guys a little bit
- 5:58more of. And this provides a tremendous
- 6:00amount of interpretability,
- 6:01accountability, and then steerability,
- 6:03which is really important that I think
- 6:04most people sleep on.
- 6:06After it's thought about things and
- 6:07basically wrote its own mini plan, it's
- 6:09time to actually act, right? And so
- 6:12here's where it'll call tools. It'll
- 6:14edit the files that it decided to uh do
- 6:16so earlier in the plan. Or maybe it'll
- 6:18run a command using command line
- 6:20interfaces, CLIs.
- 6:23After the action step is done, what it
- 6:25does is it gets the result of the tool
- 6:28call, and then it feeds all of that
- 6:30stuff back in to the observe step. So
- 6:32now we're basically running through that
- 6:34loop again, just with a little bit more
- 6:36context. And so what occurs essentially
- 6:39is we just tend to grow bigger and
- 6:41bigger and bigger and bigger. If our
- 6:43initial context was a certain size, our
- 6:46you know, second loop, it's a little bit
- 6:48bigger. Our third loop, it's a little
- 6:49bit bigger. And fourth loop and so on
- 6:51and so forth. And what this is doing is
- 6:53this is basically stacking uh more and
- 6:55more tokens into the context that the
- 6:58model can then use to plan its next
- 6:59step.
- 7:01What occurs after you go through this
- 7:02loop, you know, usually three or four
- 7:04times, is eventually the model reaches a
- 7:06point called the definition of done.
- 7:11And what the definition of done is,
- 7:13which I think a lot of people leave out
- 7:15of their agent prompts, which is
- 7:16probably why they're always underwhelmed
- 7:17by what happens, is it's the series of
- 7:19constraints and technical specifications
- 7:22required for the model to conclude that
- 7:26it no longer needs to do this loop.
- 7:29Once it reaches this definition of done,
- 7:31okay, over and over and over and over
- 7:32again,
- 7:33it notices and then it changes routes.
- 7:36So now it goes to the task complete
- 7:38route, where it generates a quick little
- 7:40final response for the user. Usually
- 7:42involves a nicely formatted answer, as
- 7:44I'm sure you guys know. Hey Nick, just
- 7:46finished your new thumbnail app build.
- 7:50And before outputting it in a window,
- 7:52either in antigravity or Codex or maybe
- 7:55Claude code, in a packaged way that you
- 7:57guys are familiar with.
- 7:59And so, obviously, if you have any
- 8:01intuition about how AI works at this
- 8:04point, if you've ever communicated with
- 8:06ChatGPT or, you know, Claude or some
- 8:09other sort of desktop AI that's nestled
- 8:11into another application that you guys
- 8:12use, you'll probably know some of this
- 8:14stuff um just as like the foundation.
- 8:16But I wanted to make it really explicit
- 8:18at the beginning of this course because
- 8:20we're going to return to each of these
- 8:21steps over and over and over again. And
- 8:23it turns out that you can heavily
- 8:25optimize all three of these. You can
- 8:27optimize the hell out of the observe
- 8:28step. You can optimize the hell out of
- 8:31the think step. And understandably, you
- 8:33can optimize the hell out of the act
- 8:34step as well. That's what we're going to
- 8:36learn. Another point I'm going to make
- 8:37in this course is that AI agents aren't
- 8:40just the large language models
- 8:41themselves. You know, I think neural
- 8:43networks and transformers are obviously
- 8:46super inherently interesting because
- 8:48they're these massive statistical things
- 8:50and these beings that can that can do
- 8:52things. They can reason. They're very
- 8:54far removed from traditional computer
- 8:55programs just 5 or 10 years ago. So a
- 8:58lot of interest goes to the LLM. But I
- 9:00want you guys to know that the LLM
- 9:02really is just a very small part of what
- 9:04most people consider AI agents these
- 9:06days.
- 9:06The LLM is of course your reasoning
- 9:08engine, right? Of course it understands
- 9:11language and of course it makes
- 9:12decisions. But it's kind of like a human
- 9:14being from like 20,000 years ago with
- 9:17like a spear in its hands, right?
- 9:19Without all of the infrastructure around
- 9:21human beings, without like your your
- 9:23house and your fireplace and your hearth
- 9:25and a place to sleep at the end of the
- 9:26night and a a society where people farm
- 9:29and produce resources and you have cars
- 9:32that you can get in and traverse a lot
- 9:33of distance. Without all the tools and
- 9:34the architecture around the
- 9:36intelligence, the intelligence is
- 9:37actually quite limited in what it can
- 9:39do.
- 9:40And that's where the rest of these
- 9:41sections come into play. So, tools, much
- 9:44like human beings, have the ability to
- 9:46read files, run code, search the web,
- 9:49call APIs, and edit files, okay? So,
- 9:52too, does this AI agent. Much like human
- 9:54beings have the ability to set a
- 9:56high-level goal and keep going until
- 9:59that task or goal is reached, you know,
- 10:01so, too, can agents. And much like human
- 10:03beings have some sort of persistent
- 10:05memory where we can keep track of things
- 10:07that we've done and then realize that
- 10:09some of those things didn't work, so we
- 10:11got to take a slightly different tack
- 10:12the next time, so, too, agents have
- 10:14things like agents.md, claw.md,
- 10:16gemini.md, access to their conversation
- 10:19history, access to auto memory files,
- 10:21and skills. And so, it's not actually
- 10:23just the LLM, for instance, that makes
- 10:26an agent work. It's really all of these
- 10:28things multiplied by the fact that, you
- 10:30know, the LLM provides us like the
- 10:32ability to be a little bit flexible. And
- 10:33that's the really big different from
- 10:35just, you know, a chatbot and then an AI
- 10:37agent. A chatbot might just be the LLM,
- 10:40okay? But, an agent takes that that LLM
- 10:43and then it adds on tools, a reasoning
- 10:45loop, memory, and so on, and so on, and
- 10:47so forth. So, as a brief example, I'll
- 10:49use an agent coding platform called
- 10:51Codex. And down here, I have a simple
- 10:54prompt where basically, I just want this
- 10:56to do a bunch of research for me on
- 10:57creatine supplementation in men.
- 11:00And what I'm doing is I'm giving it a
- 11:01brief definition of done where I'm
- 11:03saying it once you've compiled 10 plus
- 11:05empirical sources, return a structured
- 11:07report. And I'm doing this cuz I want to
- 11:08demonstrate this loop to you. And so,
- 11:10there are a bunch of other things that
- 11:11are popping up here. We have the actual
- 11:14chat window up at the top, we have its
- 11:15response, but you'll notice that in
- 11:17between, we have this sort of like
- 11:18grayed-out section here. Okay, in this
- 11:20grayed-out section is the thinking that
- 11:22the model is doing before it gets back
- 11:24to us. And so, basically, you know, if
- 11:27this was ChatGPT back from 2022 or so,
- 11:30all we would have gotten is this. But
- 11:32because I'm telling it to take actions
- 11:34in the real world, it's capable of one,
- 11:37observing. And so, it observes all of
- 11:40this text and all of its reply as
- 11:42context.
- 11:44Two, thinking. So, it's capable of doing
- 11:46a bunch of thinking on what to do next.
- 11:49And then three, acting. And so then it's
- 11:51capable of saying, "Hmm, the user
- 11:52probably wants me to do some research. I
- 11:54have access to a few tools available.
- 11:56One of the tools lets me search the web.
- 11:58Let me pump in a search term." It then
- 12:00compiled all of this information, and
- 12:02then it just repeated the same thing. It
- 12:04then with all this context said, "Okay,
- 12:06I'm observing. Not only do I have these
- 12:07messages, but I also now have a bunch of
- 12:09research. Let me think about what to do
- 12:10next. Have I achieved the goal of the
- 12:13user compiling 10 plus empirical
- 12:14sources?" And you know, after it's made
- 12:16its sort of observation and thought on
- 12:18the reasoned about it, then it's
- 12:20deciding to act. And what it's ended up
- 12:21doing after 58 seconds is giving me this
- 12:23structured evidence report. So, this is
- 12:25an example of something that might have
- 12:26looped two times, three times, but the
- 12:28more intelligent and capable these
- 12:30models are getting, um the longer that
- 12:32they're running autonomously without us.
- 12:34Hopefully, this isn't rocket science to
- 12:35anybody here, but in a nutshell, this is
- 12:37more or less what's always occurring
- 12:39non-stop every time you talk to a model.
- 12:41With all that being said, let's really
- 12:43quickly cover how to set these different
- 12:44models up. I'm going to be using Codex,
- 12:47Claude Code, and Antigravity. You don't
- 12:50need to know anything about any of these
- 12:52platforms in order to run these
- 12:53examples. And if you're already very
- 12:55familiar with, let's say, I don't know,
- 12:56Claude Code, and you've chosen to use
- 12:58that as your main agentic coding
- 12:59platform moving forward, you can skip
- 13:01over to the next section of the video.
- 13:03But I want to make sure that we all have
- 13:04an equal playing ground here, and we all
- 13:06understand how each of these platforms
- 13:07work under the hood. So, there are three
- 13:09major platforms. The first is Codex,
- 13:11which is owned, managed, and run by
- 13:13OpenAI. The second is Claude Code, which
- 13:16is owned, managed, and run by Anthropic.
- 13:18And the third is Google's Antigravity,
- 13:20which as I'm sure you can imagine is
- 13:22owned, managed, and run by Google. In
- 13:24order to start with Codex, what you
- 13:26first have to do is sign up to an Open
- 13:28AI account. The way you do so is just
- 13:30look up Open AI on Google, get to a page
- 13:33that looks anything like this, and then
- 13:35just go to the top right-hand corner
- 13:36where it says try Chat GPT. After that,
- 13:38you'll be taken to a page that looks
- 13:39something like this. You can continue
- 13:41with Google, your phone, or whatever you
- 13:43want. And if you choose to chat with the
- 13:45model and then come back at any point in
- 13:47time, just head to the top right-hand
- 13:48corner for that model again. So, I'm
- 13:50going to pretend that I haven't made an
- 13:52account before and I'll continue with
- 13:53Google. After some brief onboarding
- 13:55instructions, you'll have access to a
- 13:56page like this. But, this is just Chat
- 13:59GPT, which is more akin to a chatbot
- 14:01than anything else. We want to take this
- 14:03to the AI agent world. And so, in order
- 14:05to do that, we need to use their
- 14:07dedicated AI agentic coding platform
- 14:09Codex. So, Googling Open AI's Codex or
- 14:12something like that will take you to a
- 14:13page that looks like this, and then you
- 14:15can just click download for macOS. By
- 14:17the way, I'm on a Mac, so that button's
- 14:19automatically going to pop up for me.
- 14:21But, the Codex app is now also available
- 14:23on Windows starting March 2024th and
- 14:25beyond. The way you install things on a
- 14:27Mac is you just take this window, drag
- 14:29Codex over to applications, and then
- 14:31you're done. Once you're inside, if you
- 14:32wanted to build a website or something,
- 14:34just head over to this middle, create a
- 14:36new folder, call it whatever you want.
- 14:38So, I'll just go to downloads and then
- 14:39go a new folder, example.
- 14:42Open it within it, and now you're inside
- 14:44of this folder. Here, you can ask the
- 14:45model to do whatever you want. And so,
- 14:47what I'm going to say is make a brief
- 14:48portfolio site about Nick Surive. Keep
- 14:51it super simple and minimal.
- 14:53It'll now do some thinking.
- 14:55In our case, I actually have a design
- 14:56taste front-end skill, which improves
- 14:58its ability to create like sleek,
- 15:00high-quality looking designs.
- 15:02And now, it's looking through my own
- 15:04workspace to put together this cool,
- 15:05sexy site for me. I'm also going to ask
- 15:08it to open it.
- 15:09Uh and the way that all AI agent
- 15:11platforms work now is you have the
- 15:13ability to put a
- 15:15queued message in, which you can also
- 15:17choose to send immediately via steer. In
- 15:20In case, I'll just wait until it's done.
- 15:21It'll consume this open it message and
- 15:23then it'll just open it for me in a new
- 15:25tab. Once it's done, the open it message
- 15:27will be fed in and it's just going to
- 15:28open this for me in a new tab. Now, I'm
- 15:31kind of zoomed in here, so if I zoom in
- 15:33a little bit more, you'll see that this
- 15:34is just a a simple one-page site that
- 15:36says Nick Sarif builds clear modern
- 15:38digital work. Here's some information
- 15:40about me and here's a contact page. Not
- 15:42rocket science, but this is how easy it
- 15:43is to like build web stuff. Claude is
- 15:45pretty similar. Just Google Claude sign
- 15:47up or something like that and you'll be
- 15:49taken to a page that looks like this.
- 15:51Here, you just enter your email address
- 15:52or in my case, continue with Google. In
- 15:54Claude's case, in order to use Claude
- 15:56code, you do have to pay for it. And so,
- 15:58there is a pro plan here that's $17 per
- 16:01month with an annual subscription or 20
- 16:03bucks if billed monthly.
- 16:05I'm not working for Claude or anything
- 16:07like that. I don't have any sort of
- 16:09affiliation with Anthropic in that way,
- 16:11but I will say that I received probably
- 16:14a 100 to 200 x return on my investment
- 16:17with an agent coding platform, whether
- 16:19it's Claude or whether it's Gemini or
- 16:21whether it's Codex. So, my
- 16:23recommendation for you, if this seems a
- 16:24little bit steep, is bite the bullet,
- 16:26pay it and learn whatever you can to
- 16:29make a return on investment with that
- 16:30money in the first month because this
- 16:32stuff is really quite powerful. Assuming
- 16:34you're done, just type Claude code
- 16:36desktop download or something like that.
- 16:38You'll be taken to a page that looks
- 16:39like this, which allow you to download
- 16:41it for Mac OS, Windows or even Windows
- 16:43ARM 64. So, I'm going to give my Mac OS
- 16:46thing a quick click. Then I'll go to the
- 16:48top right-hand corner. I'll just open
- 16:49Claude up just like I did with Codex.
- 16:51That'll take me to a page like this and
- 16:53then I just drag this over to the right.
- 16:54And then once you're done, you'll be
- 16:55taken to a chat page that looks
- 16:57something like this. What we really want
- 16:59is we want this code button, so I'm
- 17:00going to give that a click. Then here,
- 17:02all we need to do is just choose a
- 17:04folder to work in and then we can put in
- 17:06a quick request. So, I'm just going to
- 17:07choose a general folder Nick Sarif. Then
- 17:10I'm going to say bypass permissions,
- 17:12which might seem a little bit scary to
- 17:13you, but it just makes the model act
- 17:15independently. Then finally, I'm going
- 17:17to say, "Hey, make a brief portfolio
- 17:19site about Nick Sheraif. Super simple
- 17:21and minimal."
- 17:22And so, just like Codex designed it a
- 17:24moment ago with its various UX uh
- 17:27features, we have the same thing here
- 17:28with Claude Code. It's going to ask to
- 17:30access some files in my folder.
- 17:33And in addition to having the message
- 17:35box, we also have this sort of grayed
- 17:36out shining uh decal here, which is sort
- 17:39of it's like sinking, if you think about
- 17:41it, as well as its tool calls.
- 17:43And what it's going to do now is
- 17:44actually build me a brief little site.
- 17:46And then just like I did before, I'll
- 17:48just say, "Open it."
- 17:50That's going to queue it, and now I can
- 17:52have a conversation with Claude. And now
- 17:54we have the actual portfolio, which as
- 17:55you guys can see here is done in
- 17:57significantly more minimal fashion,
- 17:59okay? So, this is Nick Sheraif, builder
- 18:01automation expert software engineer.
- 18:02Now, unlike with ChatGPT and then Claude
- 18:06for Anti-Gravity, odds are you probably
- 18:08already have like a Google or a Gmail
- 18:10account set up. So, all you have to do
- 18:11is just look up Google Anti-Gravity
- 18:13download, then click download for Mac
- 18:15OS. In my case, I have Apple silicon on
- 18:18Mac. If you guys don't know what you
- 18:20have, just type about this Mac, and then
- 18:22if it says Intel up here and chip,
- 18:23you're an Intel. If it's a M something,
- 18:26then you're Apple silicon. And you can
- 18:27do something similar for Windows and
- 18:29Linux, as well. And once I give that a
- 18:30click, we'll be taken to a very
- 18:32similar-looking page here, and then I
- 18:33can just drag Anti-Gravity over to
- 18:35applications. The very first time you
- 18:36open up Anti-Gravity, it'll look
- 18:38something like this. In your case, maybe
- 18:39it'll be dark mode, or maybe it'll be
- 18:41entirely light. I just have some styling
- 18:43settings, which is why mine might look a
- 18:45little different from yours. You may
- 18:46also have to log in, unless Google
- 18:48logged you in automatically. In my case,
- 18:50it logged me in automatically because
- 18:51I've used it before. Assuming that
- 18:53you've done that though, on the
- 18:54right-hand side, you'll see an agent
- 18:56model. And this agent model is very
- 18:58similar to what we saw with Codex and
- 18:59then Claude Code. All we have to do is
- 19:01just ask it to make a brief portfolio
- 19:03site about Nick Sheraif. You'll see here
- 19:04that the UX is just a little bit
- 19:06different, right? We have a little
- 19:07generating tab down here. Obviously, we
- 19:09have uh multiple settings with fast and
- 19:12Gemini 3.1 Pro. We have this little
- 19:14thinking tab. Uh it tells you how long
- 19:16it's been doing it. If it has to do any
- 19:18web searches, it does so over here.
- 19:20Hopefully you guys are seeing these are
- 19:22all just flavors that are slightly
- 19:24different, but ultimately are the same
- 19:26thing. I'm just going to write open it.
- 19:27That'll be added as a pending message,
- 19:29and then it'll open this up in a browser
- 19:31tab. As you see here, Gemini produced
- 19:33what I would probably consider to be the
- 19:34sexiest of all websites, which makes
- 19:36sense. Uh one thing I'll talk about in a
- 19:38moment is how much better it is at
- 19:40front-end design and so on and so forth.
- 19:42And yeah, we have a very simple and and
- 19:43straightforward site here. So, um this
- 19:45links to all of my resources, left
- 19:47click, YouTube, and so on and so forth.
- 19:49I probably like this one the best. From
- 19:51here on out, most of the conversations
- 19:53and the user experiences are going to be
- 19:55really similar between the agent coding
- 19:57platforms. So, while I am going to use
- 19:59multiple just to show you guys how some
- 20:01of their quirks interact, uh for the
- 20:03most part, I want you guys to know that
- 20:04the UXs are are very very similar these
- 20:07days. Like the thinking tabs, they're
- 20:08going to be the same. Some people will
- 20:10probably say that there are slight
- 20:11differences between them and so on and
- 20:13so forth. For instance, I'm a big fan of
- 20:15the little Space Invader icon in that
- 20:16Claude Code has. Uh but for all intents
- 20:18and purposes, I'm just going to assume
- 20:20that you're picking up the UX here as
- 20:22you use these models, and focus less on
- 20:24like the tiny little stuff and more on
- 20:26how to orchestrate and then prompt these
- 20:28for higher quality responses. If you
- 20:29guys want to see like step-by-step
- 20:31walk-throughs of these platforms, I'm
- 20:33going to put some little links up above
- 20:35my left shoulder here, and you can uh
- 20:37click on them anytime to go learn that
- 20:38sort of stuff. Next up, I want to talk
- 20:40about what makes these AI coding
- 20:41platforms different from one another.
- 20:44Not on a user experience um angle, but
- 20:47from an intelligence angle, from a what
- 20:49they could do angle as well. So, as you
- 20:51saw there, there were three different
- 20:53models. There was Claude, which was
- 20:55wrapped around Claude Code, Gemini,
- 20:58which was wrapped around antigravity,
- 21:00and then GPT, in my case 5.4, which is
- 21:03wrapped around Codex.
- 21:05And I think that each of these models
- 21:07are really similar at this point in
- 21:08intelligence-wise, but there are some
- 21:10pros and cons to each that basically
- 21:13like improve how they perform by a few
- 21:15percentage points. So, Claude might be,
- 21:17you know, 2% better at these, you know,
- 21:19Gemini might be 5% better at these, GPT
- 21:22might be 1% better than these. I'm just
- 21:24pulling out numbers out of my butt. But,
- 21:26I'm making them really small because I
- 21:27do want to really drive home the point
- 21:29that these models are so gosh darn
- 21:31intelligent these days that these minor
- 21:33differences only make sense at the
- 21:34bleeding edge and at the frontier. For
- 21:36most purposes, either of these are going
- 21:39to be sufficient. So, Claude has the
- 21:41most interpretable reasoning. You
- 21:43remember how I could click open that
- 21:45little reasoning tab a moment ago? Well,
- 21:47at least as of the time of this
- 21:48recording, Claude is incredible at
- 21:50making that reasoning tab really, really
- 21:52interpretable. You know exactly what
- 21:54Claude is doing at basically every step
- 21:55of the process when you use Claude code
- 21:58to visualize that reasoning. And that
- 21:59makes it really good for orchestration
- 22:02and then agentic workflows because you
- 22:05can see the decisions of the model is
- 22:06making in real time. And in doing so,
- 22:08you can also steer the model, stop the
- 22:10model, pause it, or give it new
- 22:12resources halfway through. I can't say
- 22:14the same about both Gemini and GPT. I
- 22:16think they're a lot less interpretable
- 22:17and it's a lot less accountable. You
- 22:19know, Claude is sort of a partner that
- 22:21you build things with along the way,
- 22:23whereas Gemini and GPT are almost just
- 22:24like, I don't know, they're missiles.
- 22:26You set your target, you click the
- 22:27button, and then they go. Now, there are
- 22:29some cons. Claude is a little bit slower
- 22:31unless you use fast mode, which is what
- 22:34I tend to use, although keep in mind
- 22:35that'll burn a ton of credits. And then
- 22:37I find that it's weaker at front end or
- 22:39design than a model like Gemini.
- 22:41Gemini is really good at design and
- 22:43front ends. As you guys just saw a
- 22:45moment ago, Claude picked a really
- 22:46minimalistic sleek theme. Gemini did
- 22:49some upscale stuff that still looked
- 22:51sleek, clean, but had like that
- 22:53isomorphic glass. And then GPT, maybe
- 22:56because of my design taste scale or
- 22:57something else, was kind of like more
- 22:59complex and had uh a little bit clunkier
- 23:01of a design. Well, in general, I find
- 23:03that this pattern remains the same.
- 23:05Anytime I want to design a really clean
- 23:07front end, I'm going to use Gemini for
- 23:08that. It's also got superior multimodal
- 23:10abilities. That just means there's
- 23:12actual like endpoints using the Gemini
- 23:14API um where it can understand video.
- 23:17Right now, Claude and GPT both really
- 23:18struggle with this, although you can
- 23:20build custom pipelines to do that, which
- 23:21I should have showed you guys about. It
- 23:22also has the ability to use a fast
- 23:24output, which means it writes really,
- 23:26really quickly if need be, um but they
- 23:27don't have access to a dedicated fast
- 23:30mode where you could pay more money to
- 23:31use them really quick. I think it's the
- 23:33least interpretable of the models, and
- 23:34personally I find the quality is quite
- 23:36inconsistent. There's some days when
- 23:38I'll prompt it and it'll do quite
- 23:39incredible, then other days where I'll
- 23:40prompt it and it will just absolutely
- 23:42crap the bed.
- 23:43You know, at least Claude's quite
- 23:44consistent in that way, despite the fact
- 23:46that maybe it's a little bit worse at a
- 23:47few things.
- 23:48Finally, there's GPT. There's the Codex
- 23:51series of models, the 5.4 series of
- 23:53models now. These are the best at
- 23:54back-end programming. I think they're
- 23:56also the best at like um absolute
- 23:58mathematics, which probably feeds into
- 24:00that. They're really great at
- 24:01test-driven development, and you know
- 24:03how I mentioned earlier Gemini and GPT
- 24:05are more like rockets that you point at
- 24:06a at a at a place and then they go. Um
- 24:08well, these test-driven development
- 24:11approaches essentially mean you just
- 24:12outline that definition of done, and
- 24:14then it fires and just goes autonomously
- 24:16until it reaches that. There's also
- 24:18quite a big ecosystem of different apps,
- 24:19and you know, there's a lot of um
- 24:21documentation online about how to use
- 24:23various GPT workflows and stuff like
- 24:25that, because this was the first major
- 24:27player to the AI agent market. I'd give
- 24:30it sort of like a uh you know, two out
- 24:32of three on the rest of these. I think
- 24:33Claude is much better at its
- 24:35interpretability, it's much better at
- 24:36orchestration and stuff like that. But
- 24:38GPT, being a model that just came out
- 24:40quite recently, a 5.4 anyway, is
- 24:43obviously sort of like topping the
- 24:44charts right now on a lot of stuff. Just
- 24:45some caveats there, a lot of people
- 24:47treat this as like
- 24:49>> [snorts]
- 24:49>> anathema for you to claim that, you
- 24:51know, Claude is better than GPT at this
- 24:53thing, and Gemini is better than than
- 24:55Claude at that thing. The reality is, as
- 24:57I mentioned and alluded to at the
- 24:59beginning, there are very minor
- 25:00differences between these models at this
- 25:02point. All of them are basically trained
- 25:03on the entirety of the internet as is.
- 25:05And so because of this, the slight
- 25:08differences in capabilities in the model
- 25:10tend to have more to do with like when
- 25:11they were trained and how recent it is
- 25:14versus, you know, some inherent like
- 25:15cool new design technique. Really,
- 25:17they're just training these galaxy-sized
- 25:19brains on the entire internet at this
- 25:21point. So because we're talking about
- 25:22the LLM intelligences, you know, if like
- 25:24GPT was trained after Claude, GPT's
- 25:26probably going to be a little bit better
- 25:27in certain circumstances. If Gemini's
- 25:29trained after GPT, it'll be better. But
- 25:32all that stuff resets with the next
- 25:33generation. So though I am going to be
- 25:35showing you guys some cool multi-MCP
- 25:37orchestration uh techniques later on, I
- 25:39want you to know that you don't have to
- 25:40treat all this super seriously. You can
- 25:41also just pick one model and then use
- 25:43that. Okay, next up I want to chat
- 25:45agents.md and then how to build a
- 25:47self-modifying and self-correcting
- 25:49system prompt that significantly
- 25:51minimizes the number of errors that you
- 25:53get as you build things with these AI
- 25:55agents. So for the purposes of this
- 25:57demonstration, I'm going to be using
- 25:58antigravity and through it the Gemini
- 26:00series of models. When you open up
- 26:02antigravity, you have a little window
- 26:03that looks like this. Generally, I
- 26:05divide this into three panes. You have
- 26:07your explorer on the left-hand side,
- 26:08your file editor in the middle, and then
- 26:10you have your agent on the right. What
- 26:12I'm going to do for the purposes of this
- 26:13demo is I'll just click open folder and
- 26:15then I'm going to go to antigravity
- 26:17example and just open this up.
- 26:19Okay, and what I want to do here is I
- 26:20just want to show you how all of this
- 26:21stuff works to start.
- 26:23As you guys could see on the left-hand
- 26:24side, we have a file called Gemini.md.
- 26:27Now what occurs is when you talk to this
- 26:29model over here, hey, what's up?
- 26:32Basically, what's occurring is this file
- 26:35is being prepended to the very top of a
- 26:38conversation chain. And so if I open up
- 26:40this file right now, you see how it's
- 26:41empty, there's nothing in it. Well, when
- 26:43I started this conversation and said,
- 26:45"Hey, what's up?" Okay, it knows that my
- 26:47name is Nick, but it does it knows this
- 26:49because of the fact that I'm signed in
- 26:50as Nick Surave.
- 26:52Now I want you to see what happens if I
- 26:54paste in my name is Antonio Banderas,
- 26:56refer to me as such, always always also
- 26:58always sign off super kawaii desu. So,
- 27:00I'm going to go here to the top right
- 27:01hand corner and I'll say, "Hey, what's
- 27:04up?"
- 27:05And after initializing a new model,
- 27:09notice how it's now going to return
- 27:11something quite different to what we had
- 27:13a moment ago. The reason why is of
- 27:15course this gemini.md is just a
- 27:17templated structured prompt that is
- 27:20basically always inserted into the
- 27:22beginning. Okay? The same thing applies
- 27:24with Codex, the same thing applies with
- 27:26Claude Code.
- 27:27But the names of the files are a little
- 27:29bit different. So, if I was in, let's
- 27:31say, Codex for instance, I wouldn't call
- 27:32this a gemini.md, I'd call this an
- 27:34agents.md. If I was in Claude Code, I
- 27:36wouldn't call this an agents.md, I'd
- 27:38call this a Claude.md. Whatever file you
- 27:41use here doesn't really change the idea.
- 27:43The idea is that at the very top of any
- 27:45prompt, you just have this file
- 27:47prepended to it.
- 27:48The reason why this is so powerful is
- 27:50because you now have the ability to
- 27:52statically template out the same prompt
- 27:55over and over and over again on every
- 27:57independent session. This may seem like,
- 27:59well, why don't you just copy and paste
- 28:01the same thing in instead of having to
- 28:02use this elaborate file system
- 28:03structure? And the reason why is because
- 28:05what you can do is at the very beginning
- 28:07of this file, you can actually contain
- 28:10within it like a list of lessons or
- 28:12learnings from previous instances.
- 28:15Then you can build in a like a meta
- 28:16prompt structure where before a model
- 28:18signs off, before it finishes whatever
- 28:20it's doing, it always updates that file
- 28:22with more and more and more knowledge.
- 28:24In that way, okay, you can build a
- 28:26high-quality list of like memories,
- 28:28preferences, and rules, not to mention
- 28:31things to avoid, that significantly
- 28:33improves your agent's ability to operate
- 28:35over a long time scale. And just to show
- 28:37you guys what I mean, let me show you a
- 28:38diagram. In this hypothetical instance,
- 28:41we're going to be using gemini.md.
- 28:43And basically what will occur every time
- 28:44is a new session is going to start over
- 28:46here. The agent will first read
- 28:48Gemini.md.
- 28:50You'll then give it a task like, "Hey,
- 28:52build me a website that does whatever."
- 28:55Now, it'll return the website for me,
- 28:57and then I'll say, "I don't like this.
- 28:59No dark mode."
- 29:01After I give it its feedback of no dark
- 29:03mode, rather than just correcting the
- 29:05build, it'll actually write that to my
- 29:07Gemini.md for next time, which allow the
- 29:10agent to continue working with the rule
- 29:11applied. When the session ends and a new
- 29:13session starts, now the agent will read
- 29:16the Gemini MD, but the Gemini.md will
- 29:18have an additional rule placed, okay?
- 29:20This is my file over here. It'll say,
- 29:22"No dark mode."
- 29:23And that means the next time I ask it to
- 29:25build me a website or any sort of web
- 29:26property, it'll see no dark mode, and
- 29:28then it won't make that mistake again.
- 29:29This lets your knowledge accumulate over
- 29:32sessions. The first time that you use,
- 29:34you know, Gemini or Claude Code or or
- 29:36Codex or whatever, you know, you're only
- 29:38going to have, let's say, one rule or
- 29:40one preference stored. And so, the
- 29:42number of errors that the model makes,
- 29:44errors relative like your preferences,
- 29:45will be pretty high. The second time
- 29:48that you use it, though,
- 29:49the number of errors or issues that it
- 29:51makes that don't line up with your
- 29:52preferences will go down.
- 29:54The third time, they'll go down further.
- 29:57The fourth time, it'll go down further.
- 29:59And the fifth time, it'll go really,
- 30:01really low, to the point where it maybe
- 30:02it makes zero errors at all.
- 30:04You can see that um sort of
- 30:05diagrammatically over here, with when
- 30:07you start, your thing has zero rules,
- 30:08okay? As it grows longer and longer and
- 30:11longer, you're writing more and more and
- 30:13more and more rules. Um the agents get
- 30:15better and better and better at
- 30:16understanding and then um
- 30:18anticipating as well your preferences.
- 30:20So, what does this actually look like in
- 30:21practice? Well, it's not all that
- 30:23difficult, and you can just append or
- 30:25prepend this to any Gemini, Claude, or
- 30:28agent's MD, however you like. It also
- 30:30doesn't need to be this long, although I
- 30:31did want to go into a fair amount of
- 30:32detail here with you. So, you can
- 30:34absolutely just turn this into like a I
- 30:35don't know, a three or four-line
- 30:37snippet.
- 30:38Essentially, before we start any task,
- 30:40read this entire file.
- 30:41This file contains a growing rule set
- 30:43that improves over time. At session
- 30:45start, I want you to read the entire
- 30:47learned rule section before doing
- 30:48anything.
- 30:49How it works. When the user corrects you
- 30:51or you make a mistake, immediately
- 30:52append a new rule to the learned rule
- 30:54section at the bottom of this file.
- 30:57Rules are numbered sequentially and
- 30:58written as clear imperative
- 30:59instructions. The format is category
- 31:02never or always do X because Y, and then
- 31:05here's some more formatting
- 31:06instructions.
- 31:07When do you add a rule? Add a rule when
- 31:09the user explicitly corrects your
- 31:10output. When the user rejects a file
- 31:12approach or pattern. When you hit a bug
- 31:14caused by a wrong assumption or when the
- 31:16user states a preference. Okay, and then
- 31:18it'll give some examples here of
- 31:19different rules and code. Then we have
- 31:20the learned rules down here. So, what
- 31:22I'll do, just to show you guys what this
- 31:24looks like, is I'll say, "Build me
- 31:27a simple portfolio site for Nick Saraf."
- 31:30And I'm going to have it go accomplish a
- 31:31task for me. And then, I'm inherently
- 31:33and intentionally going to give it some
- 31:36instructions.
- 31:37You see, the very first thing it did was
- 31:38analyze the gemini.md. And so, now it
- 31:41actually has this entire file as context
- 31:43inside of its thread. You can't see that
- 31:46context here because obviously they
- 31:47don't want to just muck up your your
- 31:49conversation thread, but it is literally
- 31:51like if you just pasted this entire
- 31:52thing directly in, okay? So, it's going
- 31:55to be reading that constantly as it's
- 31:56building up the rest of our website.
- 31:58And you can see that it's like it's
- 31:59built some cool terminal display here.
- 32:02It's using a library called Vit, which
- 32:03is probably like the best front-end
- 32:05library. Let's see what it does. Okay,
- 32:07this website is looking really, really
- 32:09sexy, super clean, and it clearly went
- 32:10above and beyond with my spec. However,
- 32:13I don't like how it's dark mode. So,
- 32:14what I'm going to do is go back here and
- 32:16then give it some instructions. "Quit
- 32:18doing things in dark mode."
- 32:21And the idea here is, when I give it an
- 32:23instruction like quit doing things in
- 32:25dark mode, what it's going to do is it's
- 32:26going to take my message and then say,
- 32:29"Hey, let's update our gemini.md to
- 32:32never create applications in dark mode.
- 32:34It's a user preference."
- 32:36If I scroll down here now, you can
- 32:38actually see that this style has been
- 32:40added. And so, if the next time I run a
- 32:43model
- 32:44and instantiate anti-gravity, I say,
- 32:46"Hey, I'd like you to build me a
- 32:46website." You'll actually have this up
- 32:48at the very, very top of its prompt.
- 32:51Meaning that I'm never, ever going to
- 32:52have a dark mode website again.
- 32:54In this way, this will continuously get
- 32:56closer and closer to my preferences
- 32:58until the number of rules becomes so
- 33:00exhaustive that, you know, it'd actually
- 33:01be counterproductive. In practice, I
- 33:04haven't actually hit this limit yet. I
- 33:05think this just gets better and better
- 33:06and better over time, but I could
- 33:08hypothetically see if you were to get to
- 33:09a point where there's a thousand
- 33:10independent rules, some of them would
- 33:11probably start stepping on its its toes.
- 33:14Um this sort of self-modifying Claude
- 33:16agents or Gemini.md is a very, very high
- 33:19ROI design pattern. So, whatever you're
- 33:21building with an AI agent, whether
- 33:22you're using them for business,
- 33:23personal, or programming tasks, I would
- 33:25always recommend to have something like
- 33:26this in your directory. And as you can
- 33:28see, it's now modified the site. We
- 33:29don't actually have that anymore. A lot
- 33:31cleaner, and it also fixed up the images
- 33:33and made it look really sexy. The way
- 33:34this works is at the very top level, we
- 33:36have a global Claude agents or
- 33:39Gemini.md. And these are user-wide rules
- 33:42that apply to all of the projects that
- 33:44you start. And so, the very top, you'll
- 33:46have this sort of injected, and you can
- 33:48set this using a variety of different
- 33:50formatting conventions and stuff. You
- 33:52could look it up for the specific uh
- 33:54agent platform that you're using. And if
- 33:56you're doing Claude or something like
- 33:57that, it's going to be stored in a a
- 33:59tilde. dot Claude {slash} and then there
- 34:03are a variety of other conventions
- 34:04regardless of whatever platform you're
- 34:06using that you guys can also After it's
- 34:08injected the global agents.md, it'll
- 34:11then inject the local Claude.md. And so,
- 34:14what you could do is you could have a
- 34:15global Claude.md, okay, that has
- 34:19wide-ranging user preferences updated,
- 34:21and then a local project.md that has
- 34:24specific project preferences updated.
- 34:26And then underneath, you also have uh
- 34:27skills, and then you're finally in-line
- 34:30prompt. And I'll touch on the skill
- 34:32section in a moment. But in that way you
- 34:34can collapse a ton of context and a ton
- 34:36of sort of functionality into very few
- 34:39tokens, which is important because your
- 34:40bill both per token and then the quality
- 34:42of the models tend to degrade the longer
- 34:44the token context windows get. Next up I
- 34:46want to talk a little bit about agent
- 34:48skills. And this isn't going to be an
- 34:49exhaustive resource. If you guys want a
- 34:51super in-depth way to look at skills,
- 34:53definitely just check out my full
- 34:54end-to-end Claude code skills course.
- 34:57But agent skills, for those of you guys
- 34:59that don't know, is just a simple
- 35:01repeatable way that you can standardize
- 35:04workflows.
- 35:05Now, this is important because large
- 35:07language models are very flexible. So,
- 35:09if you give them a non-super tightly
- 35:11scoped task, they'll tend to produce a
- 35:13variety of different results for you.
- 35:15Well, skills are just a way of basically
- 35:17turning that whole, you know, vagueness,
- 35:20that whole statistical variance into
- 35:22like a really straight-line
- 35:24deterministic path where it just does
- 35:25the same thing over and over and over
- 35:27and over and over again.
- 35:28And so, skills are offered now on all
- 35:30major platforms. We've all adopted them.
- 35:32So, you have Codex skills, you have
- 35:34Gemini skills, and then you also have
- 35:37Claude code skills. And they have very
- 35:39particular specs and they look really,
- 35:41really similar to one another. So, it's
- 35:42worth me at least going over to
- 35:43high-level what they look like. To make
- 35:45a long story short, these are just files
- 35:47that will exist somewhere within our
- 35:49workspace. These files will have sort of
- 35:51this little title section up here, which
- 35:53you know is a title because there'll be
- 35:54three hyphens at the top and three
- 35:56hyphens at the bottom. Inside of the
- 35:58file you can give it a name like PDF
- 35:59processing, a description like extract
- 36:01text and tables from PDFs,
- 36:04and then you can even do licenses and
- 36:05metadata and so on and so forth. I don't
- 36:07actually do any of this stuff. My skills
- 36:09are almost always just name,
- 36:10description, and then maybe some
- 36:11optional
- 36:13tools that it could use as well. Okay,
- 36:15so I just want to give you guys a couple
- 36:16of brief examples. I'm just going to go
- 36:17over to Anthropic skills because they
- 36:20have a a bunch of simple ones here that
- 36:22we can use just to gain some context.
- 36:24I'm going to go over to the skills
- 36:25folder here and then click on I don't
- 36:27know let's do algorithmic art.
- 36:29We'll go skill.md cuz that's the file
- 36:32and as you guys could see here we have
- 36:34if I click on the raw you guys will see
- 36:36we have the exact same format that I
- 36:37showed you guys earlier. So this is a
- 36:39skill that creates algorithmic art using
- 36:41a particular library and what's cool is
- 36:43it basically guides the model through
- 36:45the same thing every time to get very
- 36:47very similar algorithmic art generated.
- 36:50You can see this is a pretty long skill
- 36:51there's a lot going on right? So what
- 36:53I'm going to do is I'm just going to
- 36:54copy this whole thing and show you guys
- 36:55how this works. In this way we can copy
- 36:57and paste different standard operating
- 36:59procedures to different models and then
- 37:01get high quality results. So I'm going
- 37:03to go over here and then you know just
- 37:05because this is a one-shot prompt I'm
- 37:07just going to feed all this in
- 37:09and then I'm going to have this model
- 37:10actually create things according to the
- 37:12skill spec. So it's doing some thinking
- 37:14and now it's asking me what do we want
- 37:15to do with it and I'm going to say yes
- 37:17save as skill
- 37:19then run. And then I'm going to actually
- 37:21have this like produce some sort of cool
- 37:23algorithmic art. Now there's no template
- 37:25file or anything like that so it's
- 37:27actually going to go through the whole
- 37:27process. It's going to create both the
- 37:29skill directory which we can find right
- 37:30over here now called algorithmic art
- 37:33and then it's also going to create like
- 37:34templates and a bunch of other stuff as
- 37:36well. Okay and our algorithmic art flow
- 37:38is just finished up so I'm actually just
- 37:40going to open this so I can take a look
- 37:41at it myself.
- 37:43And we have it. There it is. This is now
- 37:45creating algorithmic art as you guys
- 37:47could see we have particles and so on
- 37:48and so forth. I'm just going to
- 37:50significantly decrease the number of
- 37:51particles
- 37:52maybe change the noise scale and the
- 37:54turbulence. Actually move this around
- 37:56and as you guys can see we we we are
- 37:57actually producing a tremendous number
- 37:59of particles here. This is this is
- 38:00actually like rendering them directly in
- 38:01my browser which is nuts.
- 38:03Um so this is indeed algorithmic art.
- 38:05It's it's really cool super sexy. I'm a
- 38:07big fan. I don't know I mean it looks
- 38:08kind of like hair but what are you going
- 38:10to do?
- 38:10I'm just going to regenerate a bunch
- 38:12maybe change the accent colors. Okay
- 38:14maybe we'll have this as my accent now
- 38:16blue and then the background will be
- 38:17kind of this and
- 38:19I don't know my cool accent will be kind
- 38:20of like this.
- 38:22There you go. That looks pretty nice.
- 38:24We can now kind of just create new ones
- 38:27as we want and then we can also just
- 38:28completely randomize them over and over
- 38:30and over and over and over again. And
- 38:31you can see it's actually still doing
- 38:33some design in the background as we go.
- 38:34So I'm just going to change the number
- 38:36of particles to really low and then I'll
- 38:38just redesign this over and over and
- 38:39over and over again.
- 38:42And I should note that like this is not
- 38:43like a you know, it's not a piece of
- 38:45software I downloaded. We actually just
- 38:46built this. It's just we built this in a
- 38:48much more standardized and you know,
- 38:50consistent way which is really cool. So
- 38:52obviously that's that's what I want. I
- 38:53want the ability to share like
- 38:55repeatable workflows where my agent can
- 38:58build things that other people have
- 38:59validated without me necessarily having
- 39:01just to like copy and paste a piece of
- 39:02software into my computer. Now remember
- 39:04earlier how I said some models are
- 39:06better at things than others and these
- 39:08few percentage point differences can
- 39:10make a lot of impact at the bleeding
- 39:12edge or the frontier. Assuming you guys
- 39:14are at the bleeding edge and the
- 39:15frontier and those percentage point
- 39:17differences stack up, then multi-agent
- 39:20MCP orchestration is the pattern for
- 39:23you.
- 39:23Basically here what happens is you let
- 39:26one model type be the manager or the
- 39:29orchestrator. And that orchestrator will
- 39:31take a task and then dole it out, okay,
- 39:34and delegate sub chunks of that task to
- 39:37different models. And so what's
- 39:38occurring here is in this hypothetical
- 39:40example we're using Claude code to be
- 39:42our manager. We then give it some task
- 39:44like, "Hey,
- 39:46make me [snorts] a SaaS app that does X,
- 39:50Y, and Z."
- 39:51And then what it's doing is it's taking
- 39:52my command and then splitting it into a
- 39:54variety of different functions. There's
- 39:56a front end task which is delegating to
- 39:58Gemini to build the UI. There's a back
- 40:01end task which is delegating to Codex to
- 40:04build the API. There'll be some testing
- 40:06that we need to occur
- 40:07that we need to do which it'll delegate
- 40:09to Codex to do the testing. Then finally
- 40:11at the end we have Claude which will
- 40:13collect and then validate the results.
- 40:15And then if there are any discrepancies
- 40:17or issues there, you know, we can loop
- 40:18that back around hypothetically to
- 40:20different models as we will.
- 40:23And so this is a little bit more of an
- 40:25advanced design pattern, and I don't
- 40:26necessarily recommend you guys sign up
- 40:28to a bajillion patterns and waste your
- 40:30tokens that way unless you have to, but
- 40:32I wanted to cover it because this is
- 40:33sort of like the next generation of
- 40:35model intelligence. It's where instead
- 40:37of just sticking with one, you're
- 40:38constantly querying different models for
- 40:40things that they're a little bit better
- 40:41at. All of this depends on this idea of
- 40:44a router.
- 40:45And so this router is more or less like
- 40:47a decision hub or like a nexus.
- 40:50When you give it a task or you give it
- 40:52some sort of input, what it'll do is
- 40:54it'll just divide it into different
- 40:56subtasks that different models are
- 40:58better than other models at. So for
- 41:00instance, if we have like a high-level
- 41:02task that has to do with replicating a
- 41:04specific SaaS app, you know, and the the
- 41:07model has decided that there's some
- 41:09footage on the internet out there that
- 41:11talks about how to build it, it'll
- 41:12actually go delegate the video watching
- 41:14step over to Gemini cuz Gemini's better
- 41:16at multimodality and their endpoints
- 41:18have built-in video understanding.
- 41:20You know, if it identifies that we need
- 41:22something with a lot of complex
- 41:23reasoning, it'll route that over to
- 41:25Claude. And if it identifies that we
- 41:26need some form of sandboxed cloud code
- 41:29execution, it'll do that in Codex cuz
- 41:31they include that built-in. And maybe,
- 41:32you know, I just wanted to show you guys
- 41:34what an example would look like if you
- 41:35had something that was outside of the
- 41:36three. If you need real-time web data,
- 41:38it might do that with Perplexity or
- 41:40Perplexity's computer or something.
- 41:42And what happens is, you know, we build
- 41:43it all by parallelizing this big sweep,
- 41:47and then at the very end we combine it
- 41:48again with this router, which is
- 41:50probably, you know, at least in my case
- 41:52almost always going to be Claude Opus
- 41:544.6, 4.7 by the time you guys are
- 41:56reading it, and then that's what
- 41:58ultimately unifies it before maybe doing
- 42:00some additional Q&A, bug fixes, and
- 42:02agent review, which I'll talk about
- 42:03later. Now all of this sounds pretty
- 42:05abstract, and you're like, "Okay, why
- 42:06don't I just have all of this done in
- 42:08one thread?" So let me show you a
- 42:09practical way to actually do it. By the
- 42:11way, all the files for this course you
- 42:12can find in the top link in the
- 42:14description below. What I'm going to do
- 42:15is go back to Claude Code and open up a
- 42:18new session. And then I'm going to
- 42:20select this folder that I've actually
- 42:21already created for this purpose called
- 42:23multi-platform orchestration.
- 42:25As mentioned, you guys will get
- 42:26everything in the description if you
- 42:28want it, and I'll also run you through
- 42:30how to create it.
- 42:31>> [gasps]
- 42:31>> But for now, what I want to do, let's
- 42:33say just hide this, is say something
- 42:35along the lines of, "Hey,
- 42:37build me a full-stack
- 42:40app that lets users enter
- 42:45a desired image to generate,
- 42:48and then it generates said image. We'll
- 42:50make this really simple because I don't
- 42:52actually want this to take forever. I'm
- 42:54kind of a time crunch today.
- 42:56And I just want you guys to see how this
- 42:58deals with that problem.
- 43:00Keep in mind in this case, Claude, which
- 43:03is the model that we're currently
- 43:04talking to, cuz it's Claude Code, is
- 43:06going to be our top-level orchestrator.
- 43:10Okay?
- 43:11Now, this is going to plan things out
- 43:13for us, which is why it's entering this
- 43:14plan mode.
- 43:16Next, what we're going to do is we're
- 43:17going to delegate all
- 43:20difficult tasks, um, like back-end tasks
- 43:23to Codex, as well as testing tasks.
- 43:27Then down at the very bottom here, you
- 43:28know, for anything related to front-end,
- 43:31we're going to delegate that to Gemini.
- 43:34And so we're going to build basically an
- 43:35ecosystem here where Claude is shuttling
- 43:37information back and forth between, uh,
- 43:40you know, Codex and Gemini for various
- 43:42things. And as you can see here, it's
- 43:43already starting to ask me, "Hey, which
- 43:45image generation API would you like to
- 43:46use?" I'm actually just going to say,
- 43:48um, Nano Banana Pro 2. It's a Google
- 43:52product.
- 43:54Okay, I'm going to submit that.
- 43:55And now what it's going to do is it's
- 43:57going to decide, "Hey, how am I going to
- 43:59delegate this work?" At the end of it,
- 44:01Claude will give me a plan, and you can
- 44:02see here that it's decided on back-end,
- 44:04front-end, and so on and so forth. And
- 44:06what it'll do now is it'll actually
- 44:08dispatch work to Gemini, Codex, and then
- 44:11itself to fix a various integration
- 44:13issues. So, I'm just going to say plan
- 44:14approved, and now it's going to start
- 44:16doing the coding. The way that Claude
- 44:17Code does this is it uses the execute
- 44:20task path for Codex. And so, what is
- 44:23occurring right now is it's just sent
- 44:25this big request in to Codex's best
- 44:27model. Okay, and now just clicking the
- 44:28button in the top right-hand corner, we
- 44:30now have a preview. And um in this case,
- 44:32Claude is now reviewing the generated
- 44:34application and doing some self-testing.
- 44:36And so, we built this image generator
- 44:38app. We've asked for a cute cat wearing
- 44:41sunglasses on a beach. This is now
- 44:43passing through to an API that Claude
- 44:46Code set up with a Gemini for the
- 44:49front-end and then Codex for the
- 44:50back-end's help. It's actually doing the
- 44:52the generation right now. And we've
- 44:53generated the cute picture of the cat on
- 44:55the beach. Looks great to me.
- 44:57The reason why you might want to do this
- 44:58is because well, it's kind of twofold.
- 45:00One, you get to parallelize your work as
- 45:02mentioned. And so, you get to build the
- 45:03front-end um using a model for which the
- 45:05front-end builder is the best. You get
- 45:07to build a back-end simultaneously using
- 45:09model by which the back-end builder is
- 45:11the best. And then you get to use an
- 45:12orchestrator, which basically eeks out a
- 45:14few percentage points increased like
- 45:16reasoning and decision-making and stuff
- 45:17like that because
- 45:19it's able to evaluate the code from both
- 45:21of these things independently without
- 45:23being polluted by the context window.
- 45:24And we're going to talk more about that
- 45:25specific review pattern later. But um
- 45:27this allows you to eke out, you know,
- 45:28more quality. The downside of this um
- 45:30prompt approach is it usually costs more
- 45:33because now you're splitting your tokens
- 45:34across multiple models just one
- 45:36provider. And usually providers will
- 45:37subsidize your token usage like Claude
- 45:40will subsidize most of its usage on the
- 45:41max plan for instance.
- 45:43Um the $200 a month that you spend on it
- 45:45is actually equivalent to like $5,000 a
- 45:47month in usage. Whereas when you build
- 45:48via API, it's usually a little bit more
- 45:50standardized. And then as a result of
- 45:51that, you end up building way more. You
- 45:53don't you don't get that cool
- 45:54subsidization. However, this is
- 45:56something that people are increasingly
- 45:57using for more complicated
- 45:58infrastructural projects, especially
- 46:00when as mentioned a minor percentage
- 46:02point or two difference in terms of
- 46:04quality is very important to you. And
- 46:06so, this is me just doing this in
- 46:07Claude, but you can obviously use, I
- 46:08don't know, Codex as the orchestrator if
- 46:10you wanted to build this in Codex. You
- 46:11could use Gemini as the orchestrator if
- 46:13you wanted to do this in, you know,
- 46:14entirely Gemini. Right now, this is the
- 46:16stack that seems to make the most sense,
- 46:18what people are talking about the most.
- 46:19If you guys are interested, the way that
- 46:20all of this stuff works under the hood
- 46:22is we basically set up a bunch of
- 46:24different servers that call Codex and
- 46:27Gemini inside of Claude. And so, that's
- 46:29why we see this using the Claude
- 46:31formatting above. It's because that
- 46:32Claude is the orchestrator that's sort
- 46:34of setting it up initially. And there's
- 46:35also a Claude.md, which describes how
- 46:37it's the manager. You know, you plan,
- 46:39reason, delegate, validate, and fix
- 46:41integration issues. When you break tasks
- 46:43down, break them into front and back end
- 46:45and test subtasks, and then delegate
- 46:47things as required. I'm going to include
- 46:48this prompt as well as everything else
- 46:50you need in order to do the same thing
- 46:51I'm down below in the description. But
- 46:53in order for this to work, you will, of
- 46:54course, need API keys for various
- 46:56platforms. And in order to get those,
- 46:57you do have to sign up to typically
- 46:58something a little bit different what we
- 47:00signed up to before. And in order to
- 47:01sign up to those, you do typically need
- 47:03to go directly to the platform, create
- 47:05an account, and then set up an API key.
- 47:07So, you can see over here, that's what
- 47:08I've done for Claude. And you can also
- 47:10do the same thing for OpenAI and then
- 47:12Gemini. Once you have those keys, you
- 47:13would just give it to whatever model you
- 47:14want to use to be the orchestrator, and
- 47:16then it would set this whole thing up
- 47:17for you, and then I'd be able to reason
- 47:19and then communicate with different
- 47:20models on your behalf. The next advanced
- 47:22prompting technique is the video to
- 47:24action pipeline. To make a long story
- 47:26short, up until quite recently, AI
- 47:28agents were forced to learn entirely
- 47:30through text descriptions of stuff. And
- 47:32the reason why is because multimodality,
- 47:34like vision, usually, at least in the
- 47:37context of video, was sort of out of
- 47:39bounds. There was just no way that we
- 47:41could feasibly take videos, which were
- 47:43millions upon millions of tokens when
- 47:45stitched together,
- 47:46you know, into some text format that an
- 47:48agent would understand.
- 47:50Well, now agents can learn from the same
- 47:51medium humans learn from. And we do so
- 47:53by combining a little bit about what I
- 47:55showed you guys earlier, okay?
- 47:56Multi-agent MCP orchestration with this
- 47:59idea of passing requests through the
- 48:02Gemini API cuz Gemini has built-in
- 48:05support for video now. Basically, uh you
- 48:07know how videos are a certain number of
- 48:09frames per second, like this video for
- 48:11instance is 30 frames a second. You can
- 48:13tell if you find a way to to slow it
- 48:15down to like
- 48:160.03.
- 48:17I'll go literally one frame every 0.03
- 48:20seconds or something like that. Well,
- 48:21what this model does is it divides
- 48:23videos into one frame per second
- 48:25instead. It then analyzes the images in
- 48:28succession and then uses a form of
- 48:31descriptive prompting to break that down
- 48:32into very, very clear steps.
- 48:35So, basically what occurs is you'll feed
- 48:36in something like a YouTube tutorial
- 48:37URL. Claude will receive the URL but
- 48:39cannot watch the video natively. So,
- 48:41instead it'll call the Gemini API.
- 48:44Gemini will watch the full video. Gemini
- 48:46will then extract the step-by-step
- 48:47instructions formatted as like a
- 48:49numbered list that's hyper-precise and
- 48:51hyper-specific.
- 48:53The structured steps will return to
- 48:54Claude via a very similar flow to what I
- 48:57showed you guys with the design. And
- 48:59then Claude will execute each using
- 49:00hyper-specific tools. Maybe if you're
- 49:02teaching somebody how to build something
- 49:03on Blender or Figma or something like
- 49:05that, you just give it access to the
- 49:06toolkit and it does it. Then the final
- 49:08result is the agent will have replicated
- 49:10the tutorial end-to-end. And in that way
- 49:12they can learn from the exact same
- 49:13medium that that we learn. So, I'll show
- 49:15you number one where I got inspiration
- 49:17from this and then number two how to do
- 49:19this for an actual task which in my case
- 49:21is going to be building a simple flow
- 49:22out in a no-code tool called N8N. So,
- 49:24first the inspiration was Spencer
- 49:26Sterling's post on X. He said he built
- 49:29an agentic system that taught itself the
- 49:31Blender donut tutorial by watching it on
- 49:33YouTube. It watched the tutorials,
- 49:35extracted the steps, filled in the gaps
- 49:37in its own tooling, and completed the
- 49:38entire thing autonomously.
- 49:40And it's quite impressive to be honest.
- 49:41Um anybody that's done any sort of 3D
- 49:43design, myself included, will know that
- 49:45like the uh way you learn how to build
- 49:47things in Blender is you watch this one
- 49:48specific tutorial that shows you how to
- 49:50build a donut. And through this process
- 49:52of building the donut, you learn about
- 49:54like textures, you learn about various
- 49:56shapes, you learn about how to modify
- 49:58them and sculpt and paint and do all
- 50:00this stuff. So, I made my own donut
- 50:01personally a few years ago. I showed it
- 50:04to all my friends, but I probably never
- 50:05touch Blender again.
- 50:07Well, the issue with knowledge like this
- 50:09is it's obviously extraordinarily
- 50:10visual, right? In order to really learn
- 50:12something, you have to watch a video.
- 50:13You can't really break all that down
- 50:15into like hyper-specific text
- 50:16instructions unless, you know, somebody
- 50:18were to just like literally go
- 50:20step-by-step. Step one, click this
- 50:22button. Step two, rotate 0.283° to the
- 50:25left. Step three, do this. So, there's a
- 50:27fair amount of nuance and flexibility
- 50:28there. And that's where video learning
- 50:30comes in handy. Human beings learn
- 50:32through video, obviously, but models
- 50:33have a tough time doing it. And so, what
- 50:35we do is we convert all of this into a
- 50:37sequence of steps. We leave some steps a
- 50:39little bit more vague, a little bit more
- 50:40general, let the model have its own kind
- 50:43of interpretability, and then give it
- 50:44some way to like screenshot its results
- 50:46to match it up to, you know, like the
- 50:48frames in the video. And so, this fellow
- 50:50here built this cool like workflow
- 50:52building studio. It's sort of like his
- 50:54own main operating system, I suppose.
- 50:56That's what this is. It's not like an
- 50:57app that he downloaded. It's something
- 50:58that he built. And then he fed in this
- 51:01along with the workflow I'm about to
- 51:02show you to have it actually like build
- 51:04the freaking thing. And it's
- 51:05communicating with this app, Blender,
- 51:07using what's called MCP, Model Context
- 51:09Protocol, which is the same thing that
- 51:10we use to communicate with the various
- 51:12models like Gemini and the Codex
- 51:14earlier. And you can get all that stuff
- 51:15in the description down below as well.
- 51:17So, I have this stored as a Claude skill
- 51:20in video to action over here. So, if I
- 51:23open this up and read the skill, you
- 51:24could see here that it actually says,
- 51:26"Extract actionable steps from YouTube
- 51:28videos using Gemini video understanding.
- 51:30Use when the user provides a YouTube
- 51:31link it wants to learn procedures,
- 51:33extract steps, understand visual
- 51:34tutorials, or turn video content into
- 51:36executable instructions." And so, what's
- 51:38occurring is it'll basically take a
- 51:39video, it'll download it for me, so then
- 51:41I'll just be able to feed in a YouTube
- 51:43URL, and then it'll convert that into
- 51:45like a highly optimized series of steps
- 51:47that, you know you would only really
- 51:48know or be able to use through the
- 51:50context of like an actual video. And so
- 51:52to demonstrate what I've done here is
- 51:54instead of using Gemini within
- 51:56Antigravity, which is sort of the usual
- 51:57design pattern, I thought I'd show you
- 51:59guys my actual stack like what I
- 52:00personally use. I think it's much easier
- 52:02if you just use the models inside of the
- 52:04tools inside of the companies that made
- 52:06them. But in my case I'm a very big fan
- 52:08of this Antigravity
- 52:10kind of container. Then inside of it I
- 52:11use Claude code. And so in a way I'm
- 52:13actually using a Google wrapper around a
- 52:15Claude code or Anthropic extension and
- 52:17that's communicating with a Claude or an
- 52:19Anthropic model. If you guys want to
- 52:21replicate the setup is as simple as just
- 52:23opening up Antigravity, heading to the
- 52:24left-hand side where it says extensions,
- 52:27downloading the Claude code for VS code
- 52:29plugin. I know it says VS code, don't be
- 52:30confused this is very similar to
- 52:32Antigravity, installing it and then you
- 52:35also have to log in here. After you're
- 52:37done you will have the exact same
- 52:38functionality that you have in the
- 52:39Claude desktop app that I just showed
- 52:41you guys earlier when we built out that
- 52:42little full stack app. Uh it's just
- 52:44you'll have it within Antigravity which
- 52:45also allows you to do things like you
- 52:47know organize your files and stuff on
- 52:48the left-hand side. So that's my
- 52:50personal stack. You don't have to use
- 52:51it. Some people judge me for it.
- 52:53Whatever, I like it, it works for me.
- 52:55Okay, so what I'm going to do is I'm
- 52:57going to find a YouTube video that I
- 52:58like and then I'm just going to feed it
- 52:59in these instructions. So I'll say I
- 53:00want you to use the video to action
- 53:03pipeline on and then I'm going to go
- 53:05grab an image. Now what I've done is
- 53:06I've found a flow that I built forever
- 53:08ago. It's a short video about 21 minutes
- 53:10that shows you how to scrape leads
- 53:11without paying for a few APIs. I'm going
- 53:13to bring that back into my Antigravity
- 53:15instance and then I'm going to do this.
- 53:18And what this is going to do is it'll
- 53:19start by invoking the skill and this is
- 53:21the UX for skill invocation. I think
- 53:23that's what it's called in English. Holy
- 53:25crap, that better be what it's called in
- 53:26English. And then it's now going to send
- 53:29that over to Gemini then receive back a
- 53:32list of highly specific instructions
- 53:34that you know understand UX I don't know
- 53:36highlight the colors of buttons and
- 53:38stuff like that and so on and so forth
- 53:39before actually running it locally on my
- 53:41computer. At the end of it, you'll get a
- 53:43super in-depth analysis that looks like
- 53:45this. So, you can actually see down over
- 53:47here, it says, "Here's the hyper
- 53:49detailed breakdown with literally every
- 53:51single step." I mean, like, "Hey,
- 53:52navigate over to this thing at 17
- 53:55seconds. Here's how to do this thing on
- 53:56that." And and so on and so on. So,
- 53:59like, it'll it'll literally it'll go
- 54:00visually as well and actually tell us
- 54:02what the end-to-end flow is going to
- 54:03look like, but then we'll also have just
- 54:05a tremendous amount of context about
- 54:06everything. Um so, what we're going to
- 54:08do now is we're going to feed that in
- 54:09and actually have this control my
- 54:10browser. So, I'm going to open up a new
- 54:12Claude code instance by clicking that
- 54:13little button above. We'll go bypass
- 54:15permissions. Then I'll say, "Use G Maps
- 54:17Scraper deep analysis.md
- 54:21to build out the same end-to-end flow
- 54:24for me." It's now going to open up a
- 54:26Chrome DevTools MCP server. It's then
- 54:28going to link that up to the end-to-end
- 54:30account. Now, it's actually thinking
- 54:31through everything that it's going to do
- 54:33using this file as a reference. And now
- 54:35it'll go through and actually control my
- 54:36browser to do the build. For simplicity,
- 54:38I'm just going to move this over to the
- 54:40right. Okay, and as we see, it just laid
- 54:42out the entire thing from left to right.
- 54:45So, it went through. It then identified
- 54:48what all of the steps were. It then
- 54:49created it inside of its own little
- 54:51conversation thread. And then it
- 54:54essentially generated what's called
- 54:55workflow JSON and then pasted it in.
- 54:57Now, this can obviously interact with my
- 54:58my browser as well. That's what it just
- 55:00did. So, it just went to the top and
- 55:01then basically imported this. What it's
- 55:03going to do now is just make some finer
- 55:05final minor changes. I'm going to
- 55:07configure the Google Sheets node and
- 55:08then we'll be on our way. So, what I'll
- 55:09do is I'll just take a screenshot of
- 55:11this and then paste it in. Then I'll
- 55:12say, "You're connected." Now, it's just
- 55:14going through and then it's selecting
- 55:15various elements. So, in this case, it's
- 55:17selecting that little search button.
- 55:18It's uh mapping the the fields and stuff
- 55:21like that. And then it'll just continue
- 55:22testing this nonstop until I have a
- 55:24working flow. You could see, you know,
- 55:25just kind of I mean, I should be moving
- 55:27this around cuz it's going to get
- 55:28confused. But you could see that it's um
- 55:30actually gone through and then pumped in
- 55:32like a specific search term. It's It's
- 55:34gone gone and basically done everything
- 55:36for me. Really, the only thing left is
- 55:37to do some sort of testing. You can see
- 55:38that uh if we actually click execute
- 55:40workflow, I'm just going to stop it here
- 55:41so I don't consume anything else. It's
- 55:43actually gone through and literally like
- 55:44scraped Google Maps for us, which is
- 55:46sweet. And it's just done so entirely by
- 55:48watching the video. So, it's entirely
- 55:49like native video understanding. And
- 55:51then it's extraordinarily detailed
- 55:53because we're we're dumping it all into
- 55:55a file, and then it can just constantly
- 55:56reference that file.
- 55:58It's then doing kind of a combination of
- 55:59like, I don't know, like ASCII or or
- 56:02text-based markup to
- 56:04uh you know, understand both the
- 56:05structure at like a micro level and then
- 56:06also like a macro level. Next, I want to
- 56:08chat this idea of stochastic multi-agent
- 56:11consensus.
- 56:12In case you guys didn't know, if you
- 56:14were to take one model, let's say Gemini
- 56:173.1 Pro High, and if you were to ask it
- 56:19like an idea question, "A, give me 10
- 56:22ideas to do X, Y, and Z."
- 56:25Every time you ask Gemini 3.1 Pro the
- 56:29same thing, it'll return a slightly
- 56:31different answer.
- 56:32Now, this property, some call it
- 56:34randomness, but I think the correct
- 56:36technical term is stochasticity,
- 56:38which is just where, due to minor
- 56:39statistical variations in the input or
- 56:42in the way that the models work, the
- 56:44output is going to be slightly different
- 56:46every time.
- 56:47The reason why this is so valuable is
- 56:49because you can exploit this tendency to
- 56:51get much, much better answers.
- 56:54For instance, let's say I run three
- 56:56times. One,
- 56:57two, and three.
- 57:00The reality is, if I run a query that at
- 57:02the very beginning says, "Give me three
- 57:05ideas for X." Okay?
- 57:08On the very first time, okay, we might
- 57:11get idea A,
- 57:13idea B,
- 57:14and idea C.
- 57:16If we were to hypothetically run this
- 57:17again, we'd probably get idea A,
- 57:20idea B,
- 57:21but just due to statistical variation,
- 57:24there is a chance that on the second
- 57:26run, it won't deliver us idea C at all.
- 57:28It'll actually deliver us idea D.
- 57:30And on the third run, maybe we do B,
- 57:33Maybe we do C, and then maybe we also do
- 57:35E.
- 57:36What stochastic multi-agent consensus
- 57:38is, you basically automate the process
- 57:40of spawning multiple agents, giving them
- 57:43slightly varied input prompts to take
- 57:45advantage of stochasticity, and then
- 57:46instead of just getting, let's say,
- 57:48three ideas, A, B, and C,
- 57:50you get to exploit stats to get all of
- 57:53the possibilities, including ones that
- 57:55might be a little rarer the model is
- 57:57less likely to actually answer with.
- 57:59And so in this way, you get A, you get
- 58:01B, you can get C, but you can also get
- 58:03D, and then you can get E. And so, you
- 58:05know, if you compare it to just one
- 58:06naive search, what we've done is we
- 58:08basically almost doubled the scope of
- 58:11the ideation. Now, mathematically, this
- 58:13is termed traversing the search space. I
- 58:15want you to pretend hypothetically that
- 58:17this like little pie chart here
- 58:19represents all possible answers to a
- 58:22question. Maybe the question is, I don't
- 58:25know,
- 58:26"What's the simplest way to get to 1
- 58:27million subscribers?" Right? This is
- 58:28something that I asked uh my my model a
- 58:30little while ago, because I'm interested
- 58:31in getting to 1 million subscribers.
- 58:33Now, obviously, I'm not just doing what
- 58:35the thing tells me, right? A lot of its
- 58:36ideas are stupid. But if you think about
- 58:38it, if I can parallelize a thousand
- 58:39agents all coming up with their own
- 58:41ideas, even if on net, the average reply
- 58:43or idea is a little bit worse than
- 58:45something I'd be able to do,
- 58:47I still get to run it a thousand times,
- 58:48right? It's like running like uh I don't
- 58:50know, like a 90 Q
- 58:52uh you know, it's like it's like
- 58:52Einstein versus 10,000 95 IQ
- 58:56researchers. It's like, well, the 10,000
- 58:5895 IQ researchers, despite lacking the
- 59:00brilliance of Einstein, they'll probably
- 59:01statistically figure it out eventually,
- 59:03right?
- 59:04So, um if this whole pie chart, to go
- 59:06back to things, is all possible
- 59:08responses, if you just run one search,
- 59:10basically what you're doing
- 59:12is you're only actually getting like a
- 59:14small chunk of all of the possibilities.
- 59:16And so instead, what we're doing is
- 59:17we're actually running multiple
- 59:18searches, you know, one search is going
- 59:20to get this, another search is going to
- 59:21get that, another search is going to get
- 59:23that, another search is going to get
- 59:25that, and and and so on and so forth.
- 59:27And then in this way, what we do next is
- 59:29we take the answers and then the replies
- 59:32of the model that should be red, and
- 59:33this one should be blue.
- 59:34And then in doing so, we get to traverse
- 59:36significantly more of that search space
- 59:38without actually necessarily consuming
- 59:40any more of our time.
- 59:41So, this is going to be kind of
- 59:42difficult to understand, and I think
- 59:44I've run out of colors here uh unless
- 59:45you've done something like this before,
- 59:48but I'll make it really simple by
- 59:49actually giving you guys a brief
- 59:50demonstration on, I don't know, some use
- 59:52case or problem that uh I think we
- 59:54probably all be able to relate to.
- 59:56Another final benefit is you get to do
- 59:58all this in parallel. So, like, you
- 59:59know, if you think about it, if you were
- 1:00:01to do one search and then do another
- 1:00:03search afterwards, and then do another
- 1:00:04search. So, for instance, let's say you
- 1:00:06have a query, "Give me three ideas for
- 1:00:07X." And then it gives you three ideas,
- 1:00:10and you're like, "Yeah, I want another
- 1:00:10three ideas." And it gives you another
- 1:00:12three ideas, and you're like, "Yeah, I
- 1:00:12want another three ideas." Well, at the
- 1:00:14end of it, you may have, I don't know,
- 1:00:15nine ideas or something, but it will
- 1:00:16have taken a certain amount of time. If
- 1:00:18the first search is 5 minutes, the
- 1:00:19second search is 5 minutes, and the
- 1:00:20third search is 5 minutes, well, you
- 1:00:22just consumed 15 minutes, right? So,
- 1:00:24instead, what this does is this just
- 1:00:26copies the idea, okay? But then it
- 1:00:29paralyzes it. So, "Hey, give me three
- 1:00:31ideas for X." And then what we do is we
- 1:00:32do one, two, and three, and in total
- 1:00:35this takes 5 minutes. Then we just
- 1:00:37combine those three answers back over
- 1:00:38here. The formal way to do stochastic
- 1:00:40multi-agent consensus, at least the way
- 1:00:41that I'm doing it here, is we'll provide
- 1:00:43a single question or prompt, then we'll
- 1:00:45do slight framing variations of every
- 1:00:47prompt that we're feeding into the
- 1:00:48model, and then we'll feed in, I don't
- 1:00:49know, I'll probably feed in like three
- 1:00:51or four or five or maybe 10
- 1:00:52simultaneously. Depends on how deep you
- 1:00:54want it to go. And then um what will
- 1:00:55happen is these will be instantiated
- 1:00:57into what are called sub-agents, okay?
- 1:00:59Which are similar to the main agent, but
- 1:01:01they operate in their own defined
- 1:01:02context window. And then all of these
- 1:01:03will just report back their answers to
- 1:01:05the parent agent. So, this parent over
- 1:01:07here is basically going to
- 1:01:09work with a whole fleet of sub-agents,
- 1:01:12and then once they're all done their
- 1:01:13work, it'll synthesize the answers. And
- 1:01:15then because what we're looking for is
- 1:01:16we're looking for like statistical
- 1:01:17variation, it'll calculate um what's
- 1:01:19called the mode, which is the frequency
- 1:01:21of each answer, and then the median,
- 1:01:22which is like the average of each
- 1:01:23answer, before ultimately combining all
- 1:01:25this to give you much better results.
- 1:01:27One final idea there is this idea of
- 1:01:28consensus. A lot of models are going to
- 1:01:31say the same things, obviously. Some
- 1:01:33models are going to say things that are
- 1:01:34quite different. And then finally, there
- 1:01:36will be outliers, which are wild cards.
- 1:01:38These wild cards here potentially
- 1:01:39brilliant, but they might only appear
- 1:01:41like 5 or 10% of the time, which is why
- 1:01:42we spawn so many of these agents that we
- 1:01:44can actually like farm these wild cards.
- 1:01:46We can we can milk them like cows. And
- 1:01:48then in that way, you can up your best
- 1:01:50ideas coming from these these fleets of
- 1:01:51agents.
- 1:01:53Um and then also save a lot of time in
- 1:01:54things like product ideation. I don't
- 1:01:56know, man, keyword search, titles for
- 1:01:58for for content, at least that's what
- 1:01:59I'm using it for, or a variety of other
- 1:02:01things. How research inventions. I'm
- 1:02:03sure Anthropic and and Google and OpenAI
- 1:02:06probably have fleets of models that are
- 1:02:07doing basically this exact same thing
- 1:02:09behind the scenes constantly. So, let me
- 1:02:10actually show you guys what this looks
- 1:02:11like in practice. I'm just going to zoom
- 1:02:13way out of this and close a bunch of
- 1:02:15these so you don't have to look at them
- 1:02:16anymore. Then I'm going to spawn a new
- 1:02:18Claude code tab over here on the right.
- 1:02:20And what I'm going to do is I'm going to
- 1:02:21use this skill that I've set up called
- 1:02:22stochastic multi-agent consensus. So,
- 1:02:24opening this up so you guys can read it.
- 1:02:26What we're doing is responding N agents,
- 1:02:29where N is just the number that you
- 1:02:30specify, with slight framing variations
- 1:02:33to independently analyze a problem, then
- 1:02:36aggregate results by consensus. We use
- 1:02:39this for decision making, ranking
- 1:02:40things, strategic analysis, or any
- 1:02:43problems where you want to filter
- 1:02:44hallucinations and surface high variance
- 1:02:46ideas.
- 1:02:48So, hypothetically, let's just say
- 1:02:50"Hey, I've struggled a lot with finding
- 1:02:52any traction on TikTok whatsoever. I've
- 1:02:54built up a bunch of accounts, and I
- 1:02:56can't seem to get more than like 1,000
- 1:02:58views per TikTok account. I'd like you
- 1:03:00to use stochastic multi-agent consensus
- 1:03:02to help me come up with possible
- 1:03:04candidate ideas to solve this. I'm going
- 1:03:06to feed this idea in, okay?" And this is
- 1:03:08a real idea, actually. We are struggling
- 1:03:10to get
- 1:03:11traction on TikTok. For whatever reason,
- 1:03:13we got 450k followers on Instagram, no
- 1:03:16problem, but, you know, the second we
- 1:03:18move things over to TikTok, we're just
- 1:03:19not really getting too many views.
- 1:03:21So, what it's going to start with is it
- 1:03:22will spawn 10 agents, all independently
- 1:03:25analyzing my TikTok problem. And every
- 1:03:27one of them will get slightly different
- 1:03:29analytical framing to maximize the
- 1:03:31diversity of ideas.
- 1:03:33Just going to zoom in here so you guys
- 1:03:34could see this, but we now have a
- 1:03:36conservative analysis. So, Nick Suriya
- 1:03:38has 287K YouTube subscribers. You know,
- 1:03:42his YouTube audience is primarily
- 1:03:43professionals. Here's a bunch of
- 1:03:45information about him. He has a small
- 1:03:47team. Here's how he's doing things and
- 1:03:49and so on and so forth. This agent over
- 1:03:51here says, "Hey, I want you to assume
- 1:03:53limited time and budget." This agent
- 1:03:55over here, "I want you to only focus on
- 1:03:57what is measurable and provable."
- 1:03:59This agent over here, you know, I want
- 1:04:01you to think about it from the end user
- 1:04:02and viewer perspective.
- 1:04:04And so, what we're doing is we're
- 1:04:04basically taking advantage of the
- 1:04:05parallelizability of models, not
- 1:04:08necessarily the base intelligence,
- 1:04:09though the intelligence is obviously
- 1:04:10important, but like we care more about
- 1:04:12like scanning and searching through
- 1:04:13space of all possible solutions really
- 1:04:15quickly. And then at the end, we're
- 1:04:17going to converge all this back with our
- 1:04:18parent agent. Now, once all these agents
- 1:04:20have turned green here, if I open up
- 1:04:22this thinking tab, you could see that
- 1:04:23it's now combining all of the
- 1:04:25information from each individual one.
- 1:04:28So, there's a bunch of suggestions
- 1:04:29saying, "Hey, you should try fresh
- 1:04:30account. You should try a device reset.
- 1:04:32You should try clean fingerprinting.
- 1:04:34Hey, you should try TikTok native hook
- 1:04:36reformatting. Hey, you should do duets
- 1:04:37with existing creators. Take advantage
- 1:04:39of the fact that you're probably bigger.
- 1:04:41Hey, you should do a series format, high
- 1:04:42posting frequency, and so on and so
- 1:04:44forth." Then you have some disagreements
- 1:04:46here as well. And these disagreements
- 1:04:47might be paid TikTok spark ads. Only one
- 1:04:49of the 10 agents suggested something.
- 1:04:51You know, in this one, they recommend
- 1:04:53using shorts, but then in this one, they
- 1:04:55recommend using a micro topic focus to
- 1:04:57build authority and audience clarity.
- 1:04:59You know, I'm not going to sit here and
- 1:05:00pretend like all these ideas are the
- 1:05:01bee's knees. Not all of them are
- 1:05:03capturing lightning in a bottle, but you
- 1:05:05run this thing long enough and you'll
- 1:05:06see eventually you will get some pretty
- 1:05:08good ideas. And the ideas will be
- 1:05:10consensus ideas, like the idea of a
- 1:05:11fresh account, but it'll also be kind of
- 1:05:13like outlier ideas with pain point
- 1:05:15framing, paid TikTok spark ads, niching
- 1:05:17down your account identity,
- 1:05:18cross-posting your Instagram Reels to
- 1:05:20YouTube Shorts first. I mean, there
- 1:05:21there there are a lot of possible ideas,
- 1:05:23right?
- 1:05:24>> [gasps]
- 1:05:24>> Now, it's opened up this consensus
- 1:05:26report, which I can visualize for you
- 1:05:28guys by clicking this button.
- 1:05:30And you can see here it's now saying,
- 1:05:31"Hey, here is the context. TikTok growth
- 1:05:33stalled at 1K views per account across
- 1:05:35multiple accounts despite this massive
- 1:05:37YouTube subs and 450,000 followers with
- 1:05:39almost 5 million Reels views a month."
- 1:05:41And then here,
- 1:05:43this orchestrator now summarizes it and
- 1:05:45says, "Hey, every agent independently
- 1:05:47identified TikTok native hook
- 1:05:48reformatting is really critical."
- 1:05:50You know, Instagram is a little bit
- 1:05:52different from TikTok hooks. Content
- 1:05:54optimized for Instagram will
- 1:05:55systematically fail TikTok's cold start
- 1:05:57test. So, you actually have to
- 1:05:58restructure it if you really want to
- 1:05:59crush. Same thing here, fresh account,
- 1:06:01clean device fingerprint. I mean, there
- 1:06:03is just so much context here, it's not
- 1:06:05even funny.
- 1:06:06And so, the reality is I would have come
- 1:06:07up with these ideas at some point, but I
- 1:06:10basically got to put, you know, a genie
- 1:06:12in a bottle and then have 500 genies
- 1:06:16simultaneously solve my wishes at 100x
- 1:06:19speed, and then aggregate all results
- 1:06:21for um, you know, I don't know, probably
- 1:06:23like three or four dollars realistically
- 1:06:25in terms of tokens.
- 1:06:27You also had a couple agents that said,
- 1:06:28"Is TikTok even worth it?" And uh, I
- 1:06:31think that's a really good question to
- 1:06:32ask because up until now, I really
- 1:06:34didn't think it was worth it. And so, in
- 1:06:35general, anytime that I recommend you
- 1:06:37have a strategic decision that you need,
- 1:06:40you can make a quick one-time trade-off
- 1:06:42of money for analysis by spawning a
- 1:06:45bunch of agents all with slight prompt
- 1:06:47variations,
- 1:06:48and then collecting the rankings
- 1:06:50reasoning to build this consensus map
- 1:06:52document. And from here, you can figure
- 1:06:54out your consensus items, your divergent
- 1:06:56items, and then your outliers.
- 1:06:59And you know, if they're consensus
- 1:07:00items, well, odds are probably because a
- 1:07:02lot of models have thought it's a good
- 1:07:03idea, you should probably do it. If
- 1:07:05there's some divergent items, well, you
- 1:07:06should probably like reason about these
- 1:07:08quite a bit before deciding whether it
- 1:07:10makes sense. And if it's like an outlier
- 1:07:12item, if there's only one out of 10
- 1:07:13agents doing it, well, it can either be
- 1:07:15a brilliant idea, in which case maybe
- 1:07:17you should give it a try, or it might
- 1:07:18just be a hallucination or some BS, in
- 1:07:20which case you don't. And so, what this
- 1:07:22allows you to do is execute with high
- 1:07:23confidence. Thank you very much, AI, for
- 1:07:25drawing that cute little That is a huge
- 1:07:28fist. That thing would be terrifying in
- 1:07:29real life. Um you know, this lets you
- 1:07:32scan a large portion of the search space
- 1:07:33in a very short period of time.
- 1:07:35And uh yeah, the actual way that you
- 1:07:37build it is very straightforward, and
- 1:07:38I'll run you guys through what all that
- 1:07:39stuff looks like I'm down below in the
- 1:07:41project description. So, just like
- 1:07:42stochastic multi-agent consensus allowed
- 1:07:45us to scan large amounts of search space
- 1:07:47in a short period of time. What we did
- 1:07:49is we independently delegated work over
- 1:07:51to agents and had them uh do things for
- 1:07:53us. So too can we take advantage of the
- 1:07:56same idea, but in my opinion get even
- 1:07:58higher quality results through this idea
- 1:08:00of agent chat rooms. What agent chat
- 1:08:03rooms are are where instead of, you
- 1:08:06know, parallelizing all the work and
- 1:08:07having all these agents try and
- 1:08:09independently solve problems, what you
- 1:08:11do is you give all of them slightly
- 1:08:12different personalities, and then you
- 1:08:14have them all debate with each other
- 1:08:16about these problems. And in doing so,
- 1:08:18they tend to deliver much higher quality
- 1:08:20responses because they're just like
- 1:08:22they're they're a little bit spikier,
- 1:08:23you know what I mean? They're not just
- 1:08:24like a generalized idea, which I'll
- 1:08:26visualize with like this interface, but
- 1:08:28you know, because they're they're
- 1:08:29butting heads with another, um
- 1:08:31eventually they ideas get really nuanced
- 1:08:33and really high quality. And so, um
- 1:08:36whether or not you visualize things in
- 1:08:37that way, that's personally how I think
- 1:08:39about things. You really get to carve
- 1:08:40out all the tiny little nooks and
- 1:08:41crannies of an idea when you debate.
- 1:08:44And so, here's a brief little
- 1:08:45visualization. We start with a problem
- 1:08:47or a prompt. We feed it in to, let's
- 1:08:50say, three agents here, agent A, agent
- 1:08:52B, and agent C.
- 1:08:53All three are given the same document
- 1:08:56called chat.json. And then what occurs
- 1:08:58is they basically cycle through a debate
- 1:09:00sequence where agent A says something,
- 1:09:02agent B says something, and agent C says
- 1:09:04something. And you know, if you do this
- 1:09:06naively, the results will probably be
- 1:09:07pretty low. But if you, I don't know,
- 1:09:09force a little bit of a spark where
- 1:09:11every agent has a slightly different
- 1:09:12opinion and they're not afraid to like
- 1:09:14state their opinion, um they'll
- 1:09:15challenge each other's assumptions. They
- 1:09:17will significantly improve the
- 1:09:19probability that you catch errors. And
- 1:09:21then this chat.json ends up being quite
- 1:09:22a valuable resource because it also
- 1:09:24shows like problem solving and stuff
- 1:09:25like that. You can then give that to an
- 1:09:27orchestrator and ultimately receive
- 1:09:29higher quality output at the end. And so
- 1:09:30it's sort of similar to what we had
- 1:09:32earlier, right? It's just instead of
- 1:09:33this operating um in parallel lanes,
- 1:09:36what these agents are doing is actually
- 1:09:38talking back and forth with each other.
- 1:09:40And so they're actually capable of
- 1:09:41having these conversations.
- 1:09:42>> [sighs and gasps]
- 1:09:43>> And I mean like I I just want you to
- 1:09:44pretend we actually spawn 10 agents.
- 1:09:46Agent one would be able to communicate
- 1:09:47with agent two, but also agent three,
- 1:09:49and also agent four, and also agent
- 1:09:51five, and also agent six. So like the
- 1:09:53total number of paths and um potential
- 1:09:57like communication,
- 1:09:58I don't really know what you want to
- 1:09:59call them, like like vectors, um goes up
- 1:10:02like crazy. And these agents,
- 1:10:04ultimately, assuming that the idea is an
- 1:10:06absolute BS, do end up at the end of it
- 1:10:08like quite quite differentiated um in
- 1:10:11their ideas and their opinions. So to
- 1:10:13show you guys what this looks like, I
- 1:10:14have another skill, which is just a
- 1:10:16repeatable workflow, to be clear, where
- 1:10:18I have this model chat. The description
- 1:10:21here is to spawn five cloud instances on
- 1:10:23a shared conversation room where they
- 1:10:24debate, disagree, and converge on
- 1:10:26solutions. They use round robin turns
- 1:10:28with parallel execution within each
- 1:10:30round for simplicity, and they trigger
- 1:10:32on the model chat multi-model debate or
- 1:10:34something else. So I have a bunch of
- 1:10:36context down over here, and you guys can
- 1:10:37grab this file for yourselves. What I'll
- 1:10:39do is I'll actually just pipe this into
- 1:10:40model chat. Okay, great. Use model chat
- 1:10:44for a similar to really work through
- 1:10:46this idea.
- 1:10:47And now it'll spark this model chat
- 1:10:50skill, which will then have them all
- 1:10:52dump shared context into a little
- 1:10:54chat.json, which I'll show you guys when
- 1:10:56it's done. Okay, so the debate has now
- 1:10:58concluded after these five agents had
- 1:11:00this conversation. Okay, we can actually
- 1:11:02see the the the chat conversation as
- 1:11:05well by going down here to this model
- 1:11:06chat. Uh, let's go latest and I'll go
- 1:11:08conversation.
- 1:11:09Um, basically what's occurred is we've
- 1:11:11given it a topic to talk about and then
- 1:11:14we've assigned a systems thinker, a
- 1:11:15pragmatist, an edge case finder, a user
- 1:11:17advocate, and then a contrarian to the
- 1:11:19task. So, first of all, the systems
- 1:11:21thinker begins, the pragmatist replies,
- 1:11:23the edge case finder goes, the user
- 1:11:24advocate goes, and so on and so forth.
- 1:11:26And you can see each of them are um,
- 1:11:28pretty pretty interestingly suggesting
- 1:11:30uh, various approaches. So, these
- 1:11:32advocates says, "Let me push back on
- 1:11:33something that challenges the consensus
- 1:11:35has glossed over, which is the clean
- 1:11:36device plus fresh account fixes seems
- 1:11:38fingerprinting is the problem. There's a
- 1:11:40separate explanation nobody has stress
- 1:11:41tested. Next content format is
- 1:11:43fundamentally mismatched to TikTok's
- 1:11:44cold start algo. And so, these are sort
- 1:11:46of arriving at similar conclusions
- 1:11:48despite the fact that uh, you know, we
- 1:11:50instantiated this separately. And then
- 1:11:52if we check out the synthesis, you can
- 1:11:53see that all of them have agreed that we
- 1:11:54need to run some diagnostics, that hook
- 1:11:56reformatting is necessary but
- 1:11:58sufficient. The high volume posting
- 1:11:59blitz two to five a day is wrong, and
- 1:12:01then fixing the IG YouTube pipeline
- 1:12:03immediately is important regardless the
- 1:12:05TikTok decision. This is something that
- 1:12:07I guess it got context out from one of
- 1:12:08my other files because um, basically
- 1:12:10despite the fact that I have 450k
- 1:12:12Instagram followers, very few of them
- 1:12:13are converting to YouTube subscribers
- 1:12:15and a lot of people, a lot of models as
- 1:12:16well, are suggesting that the reason for
- 1:12:18that is because Instagram is really
- 1:12:19blocking outbound links, which I think
- 1:12:22is actually fair. But then uh, there are
- 1:12:23a lot of, you know, disagreements as
- 1:12:25well. So, a lot of people say, "Nope,
- 1:12:26stitch duet stupid. TikTok versus IG
- 1:12:29pipeline is an either or. Device
- 1:12:30fingerprinting might not be the issue,
- 1:12:32maybe it's content mismatch, right?"
- 1:12:34And uh, there are a lot of insights that
- 1:12:36because we were able to sharpen our
- 1:12:38opinions via debate,
- 1:12:40these agents got that the previous model
- 1:12:43runs through stochastic multi-agent
- 1:12:45consensus did not. So, maybe we're
- 1:12:47looking for saves, not completions.
- 1:12:50Maybe there's just no category online
- 1:12:51yet. And although this not true. If they
- 1:12:53had the ability to research, they
- 1:12:55probably would have figured this out.
- 1:12:57Maybe it has to do with emotional
- 1:12:59moments. And then here it even gave a
- 1:13:00recommended execution plan.
- 1:13:03So as mentioned, you know, I wouldn't
- 1:13:04rely on agents for strategic advice at
- 1:13:07the moment, but I would certainly not be
- 1:13:09opposed to trading a little bit of my
- 1:13:11money for a bunch of my time back and at
- 1:13:13least ideating through the lower hanging
- 1:13:15fruit. If you run enough of these
- 1:13:17cycles, you will find pretty intriguing
- 1:13:20and interesting outlier ideas. That's
- 1:13:22just how statistics works. So you guys
- 1:13:24can get all this down below in that
- 1:13:25document. The next idea I want to talk
- 1:13:27about is this idea of sub-agent
- 1:13:29verification loops. To make a long story
- 1:13:32short, where previously we took
- 1:13:34advantage of parallelization, we're
- 1:13:36going to take a step back now to sort of
- 1:13:38serial processing.
- 1:13:40But when an agent works really hard to
- 1:13:43accomplish a task for you,
- 1:13:45it usually gets pretty biased in that it
- 1:13:49believes that its path was the best. And
- 1:13:52the reason why is because, you know, it
- 1:13:53just spent God knows how much time,
- 1:13:55energy, and compute cycles building your
- 1:13:58app or putting together your workflow or
- 1:14:00doing your taxes or whatever the hell.
- 1:14:03And because of that, you know, series of
- 1:14:05like design decisions and then issues
- 1:14:07and bug fixes, it's just very
- 1:14:09consolidated in its opinion that the way
- 1:14:11that it did what it did was the best. So
- 1:14:14if you were to ask that same agent,
- 1:14:15"Hey, can you make this better?" A lot
- 1:14:17of the time it'll look at it and be
- 1:14:18like, "Well, no, I did a pretty good
- 1:14:20job. I don't think there's any way to do
- 1:14:21it better."
- 1:14:22However, instead of just giving that
- 1:14:24agent back the entire context and
- 1:14:26saying, "Can you do it better?" a much
- 1:14:27smarter thing to do is to take all of
- 1:14:29the outputs, not the reasoning, then
- 1:14:32give the output, aka your code or your
- 1:14:34workflow or the results of your your
- 1:14:35accounting, to another agent and then
- 1:14:38say, "Hey, is this right?" Because now
- 1:14:40that second agent can evaluate purely
- 1:14:42based off output. It doesn't actually
- 1:14:44have to deal with evaluating things
- 1:14:45based off the reasoning or the intent.
- 1:14:47And so your work can end up being a lot
- 1:14:50higher quality as a result.
- 1:14:52So here's a quick example using like a
- 1:14:53coding thing where we wanted to build a
- 1:14:55rate limiter. What will happen is our
- 1:14:57first agent will implement and write the
- 1:15:00first draft of the code. This code
- 1:15:02output will pass to a reviewer agent.
- 1:15:05Now the reviewer agent is spawned with
- 1:15:07fresh context, meaning there's no tokens
- 1:15:09that are polluting its window. It has
- 1:15:11zero bias. And what it does is just like
- 1:15:13objectively speaking, you ask it, is
- 1:15:15this thing correct? Are there any issues
- 1:15:17here at first glance? Any ways you could
- 1:15:19simplify this?
- 1:15:21Now because it's treating this just like
- 1:15:22it's treating random snippet of code it
- 1:15:24finds on the internet, you know, it has
- 1:15:26no opinions. It has no inherent like
- 1:15:28desire to claim, well, this is the best
- 1:15:31way because I spent all this time,
- 1:15:32energy, and research figuring it out.
- 1:15:34And it'll be able to to look at things
- 1:15:35with, you know, those fresh eyes.
- 1:15:37From there, if it finds issues, the idea
- 1:15:39behind sub-agent verification loops is
- 1:15:41it'll list those issues and then pass
- 1:15:43the suggestions to a third agent called
- 1:15:45a resolver, which has zero context about
- 1:15:48any of this stuff as well. And so in
- 1:15:49this way an implementer, reviewer,
- 1:15:51resolver loop can get significantly
- 1:15:54higher quality results than just one
- 1:15:57agent doing everything simultaneously.
- 1:15:59If there are no issues, everything's
- 1:16:00approved, we're good to go.
- 1:16:02Otherwise, it resolves, we do some
- 1:16:04testing, and then we get the final
- 1:16:05verified code output.
- 1:16:07Are you guys noticing a trend here?
- 1:16:09Basically, all of these like advanced
- 1:16:11agent foundation advanced agent product
- 1:16:13techniques ultimately circle back to
- 1:16:16having multiple agents working in
- 1:16:17parallel. And it's really interesting
- 1:16:19because like the way that agents work
- 1:16:21themselves is they already do work in
- 1:16:22parallel. You know, a few years ago, um
- 1:16:24agents were basically just one
- 1:16:26statistical model, and you would ask the
- 1:16:28statistical model to help you complete
- 1:16:29the the the sentence or whatever, and
- 1:16:31then it would give you the most likely
- 1:16:32next token, and then it would rerun over
- 1:16:34and over and over again until it did
- 1:16:35that.
- 1:16:36Well, a few years back, um people
- 1:16:38started introducing this idea called a
- 1:16:41mixture of experts, which is instead of
- 1:16:43just having one model, what you do is
- 1:16:45you actually send the same thing to like
- 1:16:47three or four models, you average out
- 1:16:50the statistical probabilities of every
- 1:16:51word, and then you just pick what they
- 1:16:53all converged on. Very similar to what I
- 1:16:55did there with stochastic multi-agent
- 1:16:56consensus. And so this mixture of
- 1:16:58experts is sort of like the base
- 1:17:00foundation that resulted in a really big
- 1:17:02improvement in large language model
- 1:17:04accuracy, among other things like
- 1:17:06post-training and RLHF and and and stuff
- 1:17:08like that. But what's really cool is all
- 1:17:10of these frameworks basically do the
- 1:17:12same idea. You know, we we treat these
- 1:17:14mixture of experts now as themselves
- 1:17:16models, and then we prompt them with
- 1:17:19each other. We do them in parallel and
- 1:17:21then integrate their answers like
- 1:17:23stochastic multi-agent consensus. We
- 1:17:24have them debate against each other like
- 1:17:26with model chats. And now what we're
- 1:17:28doing is we're basically having them
- 1:17:29correct each other's work like with
- 1:17:31sub-agent verification loops. So all of
- 1:17:33these are just try
- 1:17:34trading off the same core foundational
- 1:17:37like features of models, which is that
- 1:17:39at the end of the day they're
- 1:17:39statistical machines. And so the more of
- 1:17:41these statistics that you can, I don't
- 1:17:43know, average out, the closer you get to
- 1:17:45the reality. Another way of thinking
- 1:17:46about this is if the implementer agent
- 1:17:48has already spent 200,000 tokens
- 1:17:50accumulating all that context, it'll
- 1:17:52literally remember every wrong turn and
- 1:17:54every dead end. It'll have a sunk cost
- 1:17:56bias. It'll say, "Well, I wrote this, so
- 1:17:57it must be right." And in a way it'll be
- 1:17:59blind to its own mistakes. But you pass
- 1:18:01it off to the super nerdy-looking
- 1:18:03reviewer agent, it has a fresh empty
- 1:18:05context. It'll only see the output, not
- 1:18:07the journey that we took to get there.
- 1:18:09No emotional attachment, although I
- 1:18:10think this is unnecessary
- 1:18:11anthropomorphization, and it'll catch
- 1:18:13what the reviewer missed. So uh let me
- 1:18:15show you guys how this actually looks
- 1:18:17like in practice. Here I have this app
- 1:18:19that I developed a while back for uh
- 1:18:21video on vibe coding, and you guys can
- 1:18:23check that out in the description if
- 1:18:24you're interested. It's where I
- 1:18:25basically put together a full end-to-end
- 1:18:27system that allowed you to um design and
- 1:18:29then syndicate a bunch of content. So,
- 1:18:31you know, this is just some app, right?
- 1:18:33This app, I don't even know if it's
- 1:18:34fully functional. Okay, no, it isn't
- 1:18:36because I had to turn it off. But,
- 1:18:37hypothetically, there's a big code base
- 1:18:39here, right? And so, what I want to do
- 1:18:40is I want to use this app to show you
- 1:18:42guys how an un
- 1:18:44biased code reviewer would take a look
- 1:18:46at the code that a previous agent had
- 1:18:48written, in this case Gemini, and and
- 1:18:50improve it. So, what I'm going to do is
- 1:18:52I'm going to go find this repo. Okay,
- 1:18:53and I found it over here. It's in the
- 1:18:54Splinter repository. Uh makes sense. I'm
- 1:18:57just going to open up a new Claude code
- 1:18:58instance.
- 1:18:59And then down over here, I'm going to
- 1:19:00say, "I'd like you to use
- 1:19:04And I just need to make sure I know what
- 1:19:05the skill is called.
- 1:19:08Agent review on the Splinter repo. It's
- 1:19:12Let's just say folder. It's in the
- 1:19:13parent folder. So, that it knows where
- 1:19:15this is. Now, that way I can still
- 1:19:17execute it within this business
- 1:19:19workspace, which I found a much better
- 1:19:21way of organizing things.
- 1:19:22And while it's doing that, I'm going to
- 1:19:23open up the skill.md.
- 1:19:25So, what the skill.md does is it spawns
- 1:19:28sub-agents to review, simplify, and
- 1:19:30verify output. It uses after completing
- 1:19:32any non-trivial implementation task, and
- 1:19:34it triggers on the words review this,
- 1:19:36agent review, self-review, or, you know,
- 1:19:38{slash} agent-review. And you can see
- 1:19:40it's already doing this. It's spun up a
- 1:19:42sub-agent called review Splinter
- 1:19:44codebase. And what this does is it
- 1:19:46reviews it for four things: correctness,
- 1:19:49edge cases, simplification, and then
- 1:19:51security. Now, like, do I know how to do
- 1:19:53all this programming under the hood? No,
- 1:19:55I don't. But, these agents certainly do.
- 1:19:57And so, we can take advantage of that by
- 1:19:58having an agent with zero context, this
- 1:20:00one here, review that entire workspace
- 1:20:03sort of independently and objectively.
- 1:20:05And now it's doing a bunch of reading,
- 1:20:07and it's going to integrate that with
- 1:20:08the suggestions of this model to give us
- 1:20:10a much higher quality output. All right,
- 1:20:12the Splinter code review just finished
- 1:20:14up, and we found 22 issues across the
- 1:20:16codebase. There's some critical ones
- 1:20:17here, some high issues here, some medium
- 1:20:20issues here, and then some low issues
- 1:20:22over there. Now, it's asking me if I
- 1:20:24want me to start fixing any of these,
- 1:20:26and I'll say, "Absolutely." And the
- 1:20:28whole idea behind this now is we're
- 1:20:30we're capable of looking at this
- 1:20:32completely objectively. You know, like I
- 1:20:34asked the initial model Gemini when I
- 1:20:36made the the app in the course like
- 1:20:37multiple times, "Hey, are there any
- 1:20:39issues here? Hey, are there any ways to
- 1:20:40make this better? Hey, what do you
- 1:20:41suspect is a problem?" And it just
- 1:20:42couldn't find it because it was so
- 1:20:43polluted by its own biases. Now, another
- 1:20:46model can. And it's very similar to like
- 1:20:49peer review in like academic um circles.
- 1:20:52It's not that like, you know, you're
- 1:20:53dumb for coming up with this code base.
- 1:20:55Like, how dare you? It's just that as
- 1:20:57you work on things more and more and
- 1:20:59more, you tend to see things a little
- 1:21:01more narrow and more narrow because
- 1:21:02you've explored a bunch of other
- 1:21:03possible paths. And the reality is
- 1:21:06the fact that you explored those paths
- 1:21:08and those don't work don't necessarily
- 1:21:09mean that if somebody else explored one
- 1:21:11of those paths, it wouldn't work either.
- 1:21:13And so, this is just a way of remaining
- 1:21:14as objective as humanly possible, which
- 1:21:15is obviously a very valuable thing to do
- 1:21:17when you're doing things like creating
- 1:21:18applications, code, um you know, sales,
- 1:21:21marketing, and all the various things
- 1:21:22that AI agents allow us to do. Next up,
- 1:21:24I want to talk a little bit about prompt
- 1:21:26contracts. For those of you guys that
- 1:21:28don't know, earlier on we chatted a
- 1:21:29little bit about a definition of done,
- 1:21:31right? Well, vague tasks, aka tasks that
- 1:21:34don't have clearly defined definitions
- 1:21:36of done, are basically the number one
- 1:21:39problem nowadays with what I would
- 1:21:42consider to be people's like disillusion
- 1:21:43with AI agents. Like, when a total
- 1:21:45novice starts using AI and then they
- 1:21:47dive into some agent to coding platform
- 1:21:49and then they just say, "Hey, build me a
- 1:21:51Netflix 2.0. Make me a million dollars.
- 1:21:54Make no mistakes."
- 1:21:55Um because of their extraordinarily
- 1:21:57poorly defined definition of done,
- 1:21:59because of the poorly defined goals,
- 1:22:01because they don't give it any
- 1:22:02constraints, because they don't give it
- 1:22:04any failure conditions,
- 1:22:06uh that model is just not going to do
- 1:22:07any any get anywhere near as high
- 1:22:09quality end-end result as if they did
- 1:22:11just follow a simple little uh
- 1:22:12step-by-step process.
- 1:22:14And so, the step-by-step process
- 1:22:15obviously you could learn,
- 1:22:17but you could also just like hard code
- 1:22:19it as a skill somewhere in your
- 1:22:20workspace or as uh you know, something
- 1:22:22in your cloud and MD and they just force
- 1:22:24your model to always have this
- 1:22:25information before you proceed.
- 1:22:27And so, for instance, if you give it a
- 1:22:28vague task like build a rate limiter,
- 1:22:30okay, it'll do pretty poorly. But, the
- 1:22:32whole idea behind a prompt contract is
- 1:22:34you basically make the user who puts in
- 1:22:36a request like this sign a mini contract
- 1:22:38and just say, "Okay, cool. The contract
- 1:22:40is, you know, here's what your goal is.
- 1:22:42Here what your constraints are. Here's
- 1:22:44what your format is and here's what your
- 1:22:45failure is. Are you good to go?" If the
- 1:22:47answer to that question is yes, now the
- 1:22:49model has actually gone through the step
- 1:22:50of defining your goal, your constraints,
- 1:22:52your format, and your failure. And so,
- 1:22:55all of your definitions are done. All of
- 1:22:57the various kind of technical spec
- 1:22:59requirements here are much more laid
- 1:23:01out. And then, the model sort of has a
- 1:23:03lot easier of a way of going about
- 1:23:04things.
- 1:23:05And so, this is very similar, if you
- 1:23:06guys are aware, to like this idea of
- 1:23:09scopes.
- 1:23:11Now, I, you know, I run like a freelance
- 1:23:13education platform, like an AI
- 1:23:15automation agency education platform.
- 1:23:17And so, scopes are a really big part of
- 1:23:18like a successful project.
- 1:23:20And so, I teach people how to define
- 1:23:21like really precise and concrete scopes,
- 1:23:24whether you're doing, you know, a small
- 1:23:25project for a client or working with
- 1:23:27some large enterprise businesses or
- 1:23:28something like that.
- 1:23:29And like a real real common issue is
- 1:23:31scopes just tend to either to be way too
- 1:23:33vague.
- 1:23:35And so, people don't actually clearly
- 1:23:36define them.
- 1:23:38Or, they end up way too restrictive. And
- 1:23:41so far that people, you know, in a in an
- 1:23:44attempt to counterbalance the vagueness,
- 1:23:45they end up going like way too specific
- 1:23:47and then the scope ends up being like so
- 1:23:48restrictive that it's like, you know,
- 1:23:50you're a slave to it and you can't
- 1:23:51change anything.
- 1:23:52And so, prompt contracts sort of help
- 1:23:53you navigate the thin line between too
- 1:23:56vague and too restrictive. And that's
- 1:23:57very similar in nature to like giving a
- 1:23:59contractor a task and then the
- 1:24:01contractor clarifying with you before
- 1:24:03they actually do the task, which I
- 1:24:05think, you know, is clearly a
- 1:24:08consequence of agents pushing all of us
- 1:24:10more towards like management style
- 1:24:11positions, where we just manage the
- 1:24:13inputs and the outputs of these things.
- 1:24:15So, I'm a big fan of defining these
- 1:24:16clearly. So, what does this actually
- 1:24:18mean in practice? Well, there's
- 1:24:20obviously a million and one different
- 1:24:21ways you can define prompt contracts.
- 1:24:23The way that I've decided to do so in
- 1:24:25this demonstration is through a skill
- 1:24:27called prompt-contract.
- 1:24:29And so basically before implementing any
- 1:24:30non-trivial task, this skill forces you
- 1:24:33to generate a structured prompt contract
- 1:24:35with goals, constraints, the format of
- 1:24:37output, and then failure.
- 1:24:39So the idea here is you're treating it
- 1:24:40just like a spec or a scope of work. Any
- 1:24:43task that produces code or some
- 1:24:45configuration settings or something like
- 1:24:46that needs to go through this process.
- 1:24:48And then this model will sort of
- 1:24:49self-analyze the request before drafting
- 1:24:52a four-section contract and then
- 1:24:54presenting it for approval. This is
- 1:24:56almost similar in nature to like the
- 1:24:58plan mode that a lot of these agent
- 1:25:00platforms now have. Like in Claude code
- 1:25:02for instance, it can enter plan mode and
- 1:25:03give you a brief little plan and have
- 1:25:05you approve the plan before it proceeds.
- 1:25:07It's just this formalizes it as a
- 1:25:08contract. And no, you're not signing
- 1:25:10your life away with Claude code when you
- 1:25:11do this. But you know, it's a simple and
- 1:25:13easy way to make sure that you get more
- 1:25:15repeatable and consistent and accurate
- 1:25:17outputs every time. So why don't I
- 1:25:18actually do this? Use prompt contracts
- 1:25:22to define this task. And then I'm just
- 1:25:24going to pretend that I'm giving it a
- 1:25:25really simple query. I'm just going to
- 1:25:27say I want you to build me a beautiful
- 1:25:30site for leftclick.ai. That's my agency.
- 1:25:34So what it's going to do is it'll begin
- 1:25:35by invoking the skill prompt contract.
- 1:25:38And I mean beautiful site is such a
- 1:25:39subjective term, right? I mean like what
- 1:25:41the heck does that even mean? And so the
- 1:25:42model is going to be essentially forced
- 1:25:44to ask me for more context on what
- 1:25:46constitutes a beautiful site to me. And
- 1:25:49in this way I will get a much higher
- 1:25:50quality site or app or whatever the hell
- 1:25:52at the end of it. Likewise, you could do
- 1:25:53this with any business task as well. It
- 1:25:55doesn't just have to be like a design
- 1:25:56task. You could set up a prompt contract
- 1:25:59for hey, email these 45 people and it
- 1:26:01could ask you like oh like what spec,
- 1:26:02you know, specifications do you do you
- 1:26:04want to confirm that they're emailed?
- 1:26:06And uh what do you want the emails to
- 1:26:07say? And what's the goal of a successful
- 1:26:09thing? And like do you have any failure
- 1:26:11parameters? If we only email 44, is that
- 1:26:13okay with you? Right? It basically
- 1:26:15forces it to be a lot more clear and
- 1:26:16then concise.
- 1:26:18So, what's happening now is it's gone
- 1:26:19through and it's actually accessed
- 1:26:20leftclick.ai. That's my current website,
- 1:26:22and then it's getting a bunch of
- 1:26:24screenshots and stuff like that. And the
- 1:26:25reason why is because it's attempting to
- 1:26:27build a context for the prompt contract.
- 1:26:29So, its first step was to analyze the
- 1:26:31request, right? What it's going to do is
- 1:26:33it'll identify what Don looks like.
- 1:26:35It'll identify some implicit
- 1:26:36assumptions. So, what am I about to
- 1:26:38force the model to assume without being
- 1:26:39told? Well, obviously an assumption is I
- 1:26:41already have a website, right? And so,
- 1:26:42it's going to go through, take pictures
- 1:26:44of my website, and see, well, if Nick
- 1:26:46wants something different from this,
- 1:26:47why? And then it's going to sort of make
- 1:26:49its own judgment to that end. And now
- 1:26:51it's actually giving me the contract.
- 1:26:52So, the goal is a single-page marketing
- 1:26:54site for Left Click. Here's some
- 1:26:56constraints. You know, we want smooth
- 1:26:58scroll animations under 500 lines of
- 1:27:00HTML. The format is this. There should
- 1:27:03be these sections. Subtle animations
- 1:27:05fade in on scroll hover states. A
- 1:27:07failure is if it looks like a generic
- 1:27:09Bootstrap template. A failure is if it's
- 1:27:11broken on mobile. A failure is if the
- 1:27:13animations are janky. The failure is if
- 1:27:15the file exceeds 500 lines. So, I
- 1:27:16actually really like this prompt
- 1:27:17contract. It's really simple and
- 1:27:18straightforward. So, I'm actually going
- 1:27:20to say go ahead and build it. But what's
- 1:27:21cool is, you know, we're now actually
- 1:27:22having a a conversation about this.
- 1:27:24We're actually agreeing on, you know,
- 1:27:26what the end result is going to be.
- 1:27:28And this is actually really similar in
- 1:27:30nature to the other thing that I want to
- 1:27:31talk to you guys about, which is um kind
- 1:27:34of related and orthogonal to prompt
- 1:27:36contracts, although it is a little bit
- 1:27:37different. And this is called a reverse
- 1:27:39prompting. Now, reverse prompting is in
- 1:27:42a similar vein, a mechanism used to
- 1:27:44clarify the quality of a prompt and
- 1:27:46improve the probability that it ends up
- 1:27:48okay. And basically the way that this
- 1:27:50works is instead of just like forcing
- 1:27:51the model to give you this contract and
- 1:27:53you having you sign off on it, it it
- 1:27:55takes it one step further, actually
- 1:27:56forces the model to ask you some
- 1:27:58clarifying questions ahead of time. So,
- 1:28:00rather than just give you a spec sheet
- 1:28:01and say, "Okay, we're good to go." what
- 1:28:02reverse prompting does is it has the
- 1:28:05model ask you a bunch of questions that
- 1:28:06you maybe didn't even think that you had
- 1:28:08to answer. The model then takes all that
- 1:28:10context and then feeds that into a
- 1:28:11prompt contract later on. Okay, so step
- 1:28:14one is when the user gives a task to an
- 1:28:15AI agent. So I don't know, this is like
- 1:28:17a website, right? Step two is the agent
- 1:28:19asks five clarifying questions back to
- 1:28:21the user before starting.
- 1:28:23Step three is when we answer and then
- 1:28:24the agent builds the correct thing on
- 1:28:26the first try. So significantly improves
- 1:28:27one-shot potential. And then if we
- 1:28:29didn't have reverse prompting, there'd
- 1:28:30be a lot of like wrong implicit
- 1:28:32assumptions here, which would result in,
- 1:28:33you know, the probability of a one-shot,
- 1:28:36which is just when the agent does it in
- 1:28:37literally one request,
- 1:28:39going down quite a bit. And so
- 1:28:41similarly, I also have a reverse prompt
- 1:28:43skill over here. And so if I go to this
- 1:28:46reverse prompt skill, you can see the
- 1:28:47way that this is set up is before
- 1:28:49implementing any non-trivial build, ask
- 1:28:51the user five dynamically generated
- 1:28:52clarifying questions to surface
- 1:28:54non-obvious preferences, assumptions,
- 1:28:55and constraints.
- 1:28:57So when to trigger before starting
- 1:28:58implementation, step one, analyze the
- 1:29:00request, figure out some stated
- 1:29:02requirements, implicit assumptions,
- 1:29:05some decision points, failure modes, and
- 1:29:06taste-dependent choices, right?
- 1:29:08And so likewise, if I instead wanted to
- 1:29:10build, let's say,
- 1:29:12something for
- 1:29:15build a beautiful site for One Second
- 1:29:17Copy, which is my old content writing
- 1:29:18company, which we just had to shut down
- 1:29:19a few days ago.
- 1:29:21Uh as you guys can imagine, content
- 1:29:22isn't super in these days.
- 1:29:25Oh, and then um use the reverse prompt
- 1:29:28skill
- 1:29:30and chain it together with prompt
- 1:29:32contracts after.
- 1:29:34What you could see is we're now engaged
- 1:29:36significantly more than we were before.
- 1:29:38Before, I just say, "Build me a
- 1:29:40beautiful site." Probability that it
- 1:29:41gets what I want right on the first try,
- 1:29:43pretty damn low. What it's doing now is
- 1:29:45it's asking a bunch of clarifying
- 1:29:47questions to confirm whether or not, you
- 1:29:49know, this site is as I want it to be.
- 1:29:52And then after I feed it back that
- 1:29:53information, it'll then take that and
- 1:29:55use that to construct essentially that
- 1:29:57that prompt contract that we had before.
- 1:29:59So here's what the conversation looks
- 1:30:00like. What's the primary goal of the
- 1:30:02site? Brand credibility, sales funnel,
- 1:30:04lead gen? You know, what I want is just
- 1:30:05brand credibility. Should it be a single
- 1:30:07static page site or should I build it in
- 1:30:09some other framework? No, I wanted a
- 1:30:10simple site. What's the vibe? You know,
- 1:30:12it's AI content writing. You know,
- 1:30:14should I do a clean modern SaaS
- 1:30:16aesthetic like linear versus L? Do I
- 1:30:18want something different? Yeah, I want
- 1:30:20like linear but white.
- 1:30:22You know, should I generate the copy
- 1:30:23from context or use some placeholder
- 1:30:25content? No, you're cool. You can
- 1:30:26generate it from here.
- 1:30:28Now, once we've clarified everything,
- 1:30:30what this model is going to do is use
- 1:30:32all this information to outline the
- 1:30:33prompt contract using the prompt
- 1:30:35contract skill. And now you can see it's
- 1:30:36invoking the skill as well.
- 1:30:38And here we have a contract. It'll be a
- 1:30:40single page static site for one second
- 1:30:41copy linear white aesthetic five
- 1:30:43sections to play ready. Here's some
- 1:30:45constraints.
- 1:30:46Here's the format. Maybe I don't like
- 1:30:47the format. Maybe I don't want it inside
- 1:30:49of active. You know, I want it somewhere
- 1:30:50else.
- 1:30:51Uh but anyway, in in this case, maybe I
- 1:30:53want to look good and and build it.
- 1:30:55Now, just to show you guys an example of
- 1:30:57how much higher quality we can get when
- 1:30:58we actually do this. This is the um
- 1:31:00website that uh it just built for us.
- 1:31:02I'm just going to refresh this puppy and
- 1:31:03take it to a new window because it gets
- 1:31:05cut off on that window. This is here we
- 1:31:07have those cool sexy animations. As we
- 1:31:09scroll down, we also have some
- 1:31:11information. Um it's it's light theme,
- 1:31:13right? We have these really minimalistic
- 1:31:15requirements here. Information about
- 1:31:17myself, some services page, words from
- 1:31:19happy clients, and then ultimately like
- 1:31:20a CTA. And so, you know, the reason why
- 1:31:23it was able to get much closer to what I
- 1:31:24wanted, which was a minimalistic white
- 1:31:26high-end aesthetic, is just because
- 1:31:27like, you know, I I had it outlined in a
- 1:31:29contract. As I'm sure you guys can
- 1:31:30imagine, you can employ the same
- 1:31:31approach for whatever the heck you want,
- 1:31:33whether you're building a site or you
- 1:31:34are you know, selling to people or you
- 1:31:36are doing some sort of bookkeeping or
- 1:31:38accounting. It's all just about uh
- 1:31:39building out a very strong definition of
- 1:31:41done. And the model can assist you with
- 1:31:43this. You don't actually have to sit
- 1:31:44down and laboriously write it all out
- 1:31:45yourself. And that takes us to the
- 1:31:47initial demo that we started with, which
- 1:31:49was the multi-agent Chrome MCP manager.
- 1:31:54Now, basically, at the very beginning of
- 1:31:56this course, you didn't understand how,
- 1:31:59you know, one agent could spawn a bunch
- 1:32:00of other agents. You didn't understand a
- 1:32:02lot of like the parallelization plays.
- 1:32:04He also didn't understand that uh you
- 1:32:06know, you could have agents actually
- 1:32:08chat with each other and communicate. He
- 1:32:10didn't understand the idea behind using
- 1:32:12one agent to verify the work of another.
- 1:32:14He didn't understand the idea behind
- 1:32:15delegating to multiple different types
- 1:32:17of models. What's really cool is the
- 1:32:19multi-agent Chrome setup that I showed
- 1:32:22you guys where we had, you know, five or
- 1:32:2310 agents all operating independently in
- 1:32:25their own browsers in their own
- 1:32:26workspaces. All of that just feeds off
- 1:32:29of this this idea or this concept um of
- 1:32:32you know, agents increasing their level
- 1:32:34of communication with other agents.
- 1:32:36And so, essentially, if you think about
- 1:32:38this logically, you know, if I were to
- 1:32:39do this uh with like a single agent. So,
- 1:32:41let's just say one agent.
- 1:32:44You know, it's not actually rocket
- 1:32:45science to have one agent use a browser
- 1:32:47these days. There are built-in skills
- 1:32:49called MCPs, model context protocols
- 1:32:51basically, that you can just pipe in and
- 1:32:53immediately connect to and it can do
- 1:32:54everything for you, okay? It can It can
- 1:32:56launch Chrome and then it can control
- 1:32:57things on the page and and whatnot. It
- 1:32:58can do that.
- 1:33:00>> [gasps]
- 1:33:00>> So, you know, the issue is it just takes
- 1:33:02a lot of time. We'll receive the target
- 1:33:04URL. We'll launch Chrome by the dev
- 1:33:06tools MCP. We'll navigate to the
- 1:33:07website. We'll take a screenshot. And
- 1:33:10you know, in my case, this over here was
- 1:33:11like um specific for me,
- 1:33:14which was just page
- 1:33:16or rather form fills.
- 1:33:19After that, we'll identify the form,
- 1:33:20extract the form fields, generate a
- 1:33:21personalized message, fill the fields,
- 1:33:23and then click submit. Um but you know,
- 1:33:25this is still something that's occurring
- 1:33:26linearly. And because of linear
- 1:33:28constraints, you know, unless you are
- 1:33:30using uh I don't know, like a a Gemini
- 1:33:32flash model or you're using fast mode
- 1:33:34and burning through your claw token uh
- 1:33:36usage limits, this is going to take a
- 1:33:37fair amount of time.
- 1:33:38This process over here, literally to
- 1:33:40just like launch the browser, could take
- 1:33:425 seconds. This process to navigate to
- 1:33:43the website could take 5 seconds. Taking
- 1:33:45a page screenshot could take 15 seconds.
- 1:33:47Identifying the contact form could take
- 1:33:49a minute. You know, if you stack it all
- 1:33:50up, basically what's occurring is this
- 1:33:52whole process here
- 1:33:54might take literally 2 to 3 minutes
- 1:33:57per form if you're operating naively
- 1:33:59using a slower model. And if you're
- 1:34:00operating non-naively, if you're using a
- 1:34:02smarter model, then obviously you have
- 1:34:04to weigh that against cost and and and
- 1:34:05token usage and stuff like that. So, I
- 1:34:07don't know. Let's hypothetically say, in
- 1:34:09my case, I wanted to reach out to
- 1:34:11you know, 1,000 people.
- 1:34:14Well, if it takes me two to three
- 1:34:15minutes a form, that's 1,000 * 2. That's
- 1:34:182,000 minutes, which divided by 60 is
- 1:34:20like 30 hours or something like that,
- 1:34:22right? That's a very long time. It's
- 1:34:24going to take me a whole day.
- 1:34:25So, instead of just doing one agent,
- 1:34:27what I'm going to do is I'm basically
- 1:34:29going to give every agent its own both
- 1:34:31Chrome instance and then even its own a
- 1:34:32workspace and then open up its autonomy
- 1:34:35so that it can make some advanced
- 1:34:36decisions to basically help it build its
- 1:34:38own tooling if it needs in order to like
- 1:34:40navigate website pages or whatever. Now,
- 1:34:41what this is going to look like is
- 1:34:42pretty similar to our previous
- 1:34:45you know, stochastic multi-agent
- 1:34:47consensus prompt. We're basically we
- 1:34:49have a user up top, okay? And this is
- 1:34:51us.
- 1:34:52And what we're going to do is we're
- 1:34:54going to give all the context about our
- 1:34:57task, whatever it is that we want, you
- 1:34:58know, fill out a form or I don't know,
- 1:35:00do some lead gen to an orchestrator
- 1:35:02agent, which in this case I'll do Claude
- 1:35:05and we'll just do Opus.
- 1:35:07Which
- 1:35:08in my case is going to be 4.6, maybe in
- 1:35:09your case it's a better model.
- 1:35:11And then what that will do is it'll
- 1:35:12spawn and set up, you know, however many
- 1:35:15agents we want in separate windows. That
- 1:35:18will then all in parallel navigate to
- 1:35:20the site, find the form, fill the
- 1:35:21fields, and then do the submission.
- 1:35:23And so, basically, instead of it taking
- 1:35:24two minutes per form, we can do is we
- 1:35:26can actually submit, you know, however
- 1:35:27many forms. So, I don't know. Let's say
- 1:35:28we have like 10 agents. We'd submit 10
- 1:35:31forms in the same amount of time it took
- 1:35:32to submit uh one. So, maybe for us it'll
- 1:35:34be 120 seconds.
- 1:35:36And then what we do is we just we just
- 1:35:37increase this as necessary. I mean, I
- 1:35:39could theoretically have 500 operating
- 1:35:40if I had the computing power. So, you
- 1:35:42know, if previously it was one form
- 1:35:45in
- 1:35:46uh what did I say? Two minutes.
- 1:35:48That means the form per minute rate is
- 1:35:50like 0.5 forms a minute, right?
- 1:35:53But now if we spin up 10, and we do 10
- 1:35:56in 2 minutes, we're up to five a minute.
- 1:35:59If we spin up, I don't know, 100,
- 1:36:01then we're up to 50 a minute.
- 1:36:03And you know, if my goal was 2,000 a day
- 1:36:05and we're at 50 a minute, then obviously
- 1:36:062,000 divided by 50 means we can get
- 1:36:08this whole thing done in 40 minutes. And
- 1:36:10you know, depending on the the the list
- 1:36:12and whatever the heck you got, obviously
- 1:36:13the constraints change. But um this is
- 1:36:15how you can have multiple Chrome
- 1:36:17instances operating simultaneously, it
- 1:36:19navigating the website and stuff like
- 1:36:21that. What I have here is a skill called
- 1:36:23multi-agent-chrome.
- 1:36:25And again, this is something you can
- 1:36:26implement using whatever um context
- 1:36:29framework you want, whether it's a
- 1:36:30skill, whether it's like a Claude Gemini
- 1:36:32or Agent 7D, whatever the heck you want.
- 1:36:34What this basically forces it to do is
- 1:36:35to orchestrate parallel browser
- 1:36:36automation using multiple Chrome
- 1:36:38DevTools MCP instances. And this is used
- 1:36:41when a task requires doing the same
- 1:36:42browser action across many targets
- 1:36:44simultaneously. So some good examples
- 1:36:46are submitting forms, filling apps,
- 1:36:47scraping pages that need JavaScript
- 1:36:49rendering, and and whatever.
- 1:36:50And so what's occurring down here is
- 1:36:52basically we have a top-level business
- 1:36:54workspace, which is sort of the folder
- 1:36:56that I'm in right now. And this actually
- 1:36:58interacts with a bunch of Chrome agents,
- 1:37:00which all have their own little MCP
- 1:37:03servers, their own little um quad.mds,
- 1:37:06and so on and so forth.
- 1:37:07And then they all communicate with a
- 1:37:09centralized chat.
- 1:37:10And if they run into problems on
- 1:37:12websites, if they have any reports they
- 1:37:13want to give, basically what happens is
- 1:37:15this orchestrator just checks the chat
- 1:37:17every 30 seconds or so, okay?
- 1:37:19So the very first step is it determines
- 1:37:20how many agents are needed, then it
- 1:37:22launches all the Chrome instances, and
- 1:37:23resets the chat file because uh you
- 1:37:25know, previous runs may have that. And
- 1:37:27then you can see how every individual
- 1:37:29sub-agent actually monitors its own um
- 1:37:31task list by basically just pumping
- 1:37:33things into a chat.
- 1:37:35This is one of the simplest and easiest
- 1:37:36ways of getting this specific design
- 1:37:38pattern done. As mentioned, you guys can
- 1:37:39get this down below if you want, but I'm
- 1:37:41just going to give you guys a simple
- 1:37:42example, which in my case is going to be
- 1:37:43just finding Vancouver rentals because
- 1:37:45I'm, you know, considering getting a
- 1:37:46rental um down there. And so, you know,
- 1:37:49rather than have it give you crappy
- 1:37:51results, this thing can actually
- 1:37:52navigate like Craigslist, Facebook
- 1:37:54Marketplace, Kijiji, whatever the heck
- 1:37:55you want. Uh and the the specific script
- 1:37:57is right over here. So, hypothetically,
- 1:37:59what it'll do is it'll just open up a
- 1:38:00new window. And then I'll go over here
- 1:38:02and then I'll write uh I want to find
- 1:38:05a rental in Vancouver, needs to be
- 1:38:0815-minute walk from the
- 1:38:11Granville SkyTrain station downtown. Use
- 1:38:15multi-agent
- 1:38:16Chrome to navigate
- 1:38:18through sites and give me
- 1:38:21high-quality sleek places
- 1:38:24under 2.5, let's say
- 1:38:272K to 2.5K.
- 1:38:29Other restrictions, like one bed, one
- 1:38:33bath.
- 1:38:34Reasonably
- 1:38:35near the water. Needs AC built in. Okay.
- 1:38:39So, I'm giving it a high-level um, you
- 1:38:42know, piece of instruction. And sorry,
- 1:38:43what I meant to do is actually do prompt
- 1:38:45contract after this. And now I want it
- 1:38:47to give me like a very clear contract.
- 1:38:49So, it's going to give me a list of five
- 1:38:50to 10 rental apartments. Why don't we
- 1:38:52say 20 rental apartments?
- 1:38:55And then I'll say 1.2 km is fine.
- 1:38:58We'll say near water south of Drake or
- 1:39:00west of Burrard.
- 1:39:02Okay, I'm just going to make some
- 1:39:03changes here. And then I'll say that
- 1:39:05sounds pretty good. Go for it. And now
- 1:39:06it's going to actually launch the
- 1:39:08multi-agent Chrome scraping. So, it's
- 1:39:10then going to invoke the skill. I'm just
- 1:39:12going to keep my hands off.
- 1:39:14What it'll do next is actually spawn um
- 1:39:16four [clears throat] parallel Chrome
- 1:39:17agents, one per rental site. So, it'll
- 1:39:19determine that there are four rental
- 1:39:20sites that it's going to be running
- 1:39:22through. And uh it'll just have one
- 1:39:24Chrome instance sort of do everything
- 1:39:25there.
- 1:39:26Uh per site. So, now we have the four
- 1:39:28instances. I'm just going to open this
- 1:39:29up here. Open this up here. I'll move
- 1:39:33this one down over here. And I'll also
- 1:39:35move this one down over here.
- 1:39:37Obviously, you could use an approach
- 1:39:38like this for pretty nefarious purposes.
- 1:39:40Um so, you do have to be cognizant of
- 1:39:42that that a lot of people and websites
- 1:39:44are probably, um, you know, they're
- 1:39:46looking to verify whether or not you are
- 1:39:47a person. And so there are multiple
- 1:39:50things you can do to get around that if
- 1:39:51you so wanted to, like using custom, um,
- 1:39:54browser fingerprinting and whatnot. And
- 1:39:56I think that's a story for another
- 1:39:57course because I don't really want this
- 1:39:59course to be accused of showing you guys
- 1:40:00how to spin up like 500 Chrome instances
- 1:40:02scraping all sorts of illicit
- 1:40:03information on the internet using unique
- 1:40:05browser fingerprints. But, uh, that
- 1:40:06stuff is definitely possible and there
- 1:40:08are probably like a lot of people doing
- 1:40:10stuff similar to this right now that are
- 1:40:11just way farther ahead in terms of
- 1:40:12their,
- 1:40:13uh, you know, understanding of agents
- 1:40:15and stuff like that. Now, after the
- 1:40:1630-second or so wait time, these will
- 1:40:18receive their instructions and they'll
- 1:40:19actually check the main thread and then
- 1:40:20they'll load in their websites. So, this
- 1:40:22one up here spawned, uh, lib. And I've
- 1:40:24never used that site before. This one's
- 1:40:26padmapper.com, which is another one. You
- 1:40:28know, these are all like websites and
- 1:40:30resources I probably would not have
- 1:40:31looked at. And as a result, I'm going to
- 1:40:33get some more of like a search spread.
- 1:40:35I'm going to, again, for a big chunk of
- 1:40:37the search space much faster than if I
- 1:40:39were to have done all this stuff
- 1:40:40manually. What's cool is these consume
- 1:40:42directly into pages for me. These can
- 1:40:44click on links and stuff like that.
- 1:40:45They're obviously modifying filters and
- 1:40:47and whatnot autonomously so that they're
- 1:40:49not just getting a bunch of bogus
- 1:40:50results. And at the end I get a
- 1:40:51high-quality filtered list of apartments
- 1:40:53that are,
- 1:40:54you know, within my specifications.
- 1:40:56Okay, I just turned my camera up because
- 1:40:57I wanted some additional room in the
- 1:40:58bottom left-hand side to really drill a
- 1:41:00few important points home. The first is
- 1:41:02your context window. Now, remember
- 1:41:05earlier how we talked about the
- 1:41:06claude.md,
- 1:41:08the gemini.md,
- 1:41:10and then the agents.md.
- 1:41:12Over here we just have Claude, but I
- 1:41:14just want you to treat this as all, uh,
- 1:41:16three of them.
- 1:41:17That's not the only thing that gets
- 1:41:19injected, so to speak, in your context.
- 1:41:24You have a variety of other things.
- 1:41:26Now, you have your system prompt up top.
- 1:41:30You then have the claude.md, agents.md,
- 1:41:32and whatever else.
- 1:41:34You have a file, at least in Claude
- 1:41:36code, called memory.md,
- 1:41:39although there are analogs in other
- 1:41:41coding platforms.
- 1:41:44And then you also have skills and tools.
- 1:41:49We've chatted a lot about MCP over the
- 1:41:52course of the last hour and a half or
- 1:41:54so, right? Well, MCP is a type of skill
- 1:41:57and tool.
- 1:41:58We also have the actual skills
- 1:42:00themselves. So, remember the agent
- 1:42:02reviewer when we were doing sub-agent
- 1:42:04verification loops?
- 1:42:06Well, that was an example of a skill.
- 1:42:09You think to the prompt contracts,
- 1:42:12those were examples of skills.
- 1:42:16And the reason why I'm going into depth
- 1:42:18here is because each of these sections
- 1:42:21can consume a tremendous number of
- 1:42:23tokens.
- 1:42:25And you're not given an unlimited number
- 1:42:26of tokens to start with. Everything in
- 1:42:28life is finite, including uh you know,
- 1:42:31your Claude or your Gemini context
- 1:42:33window.
- 1:42:34Now, most models right now are somewhere
- 1:42:36between
- 1:42:38uh I don't know if it's like 4.6, we're
- 1:42:40probably talking 200k to 1 million.
- 1:42:44If we're talking Gemini, you know, we
- 1:42:46have like 3.1 and there there are a
- 1:42:48couple other ones, obviously. But by the
- 1:42:50time you guys are watching this, they'll
- 1:42:51probably be more.
- 1:42:53And then, you know, you have uh GPT 5.4
- 1:42:57and then Codex 5.3, but the 5.4 is
- 1:42:59coming out.
- 1:43:00You know, most models nowadays have
- 1:43:03somewhere in the realm of between 200k
- 1:43:05to 1 million.
- 1:43:07And to be clear, um a token is not a
- 1:43:09word.
- 1:43:11A token is about
- 1:43:140.7
- 1:43:16words.
- 1:43:17So, if you think about it in that vein,
- 1:43:18what this means is these 200k tokens
- 1:43:21sort of actually equate to somewhere
- 1:43:23between like 140,000 words to about
- 1:43:26700,000 words, okay?
- 1:43:28Um but this context window obviously it
- 1:43:30filled up the more that you talk with
- 1:43:32it. And unfortunately, one common and
- 1:43:35major problem in large language models,
- 1:43:38specifically the types that we're
- 1:43:39dealing with in this course, are as time
- 1:43:42goes on and you talk to it more and more
- 1:43:45and more,
- 1:43:47what you find is the average quality
- 1:43:51goes down.
- 1:43:53So, quality
- 1:43:55as a factor or a byproduct of token
- 1:43:59count, typically starts pretty high up
- 1:44:01here at maybe, I don't know, 100%.
- 1:44:05And then the longer and longer and
- 1:44:07longer the number of tokens in your
- 1:44:08context, the lower and lower and lower
- 1:44:11the quality gets.
- 1:44:12So, maybe this is a 10K.
- 1:44:15Maybe this is at 50K.
- 1:44:18Maybe this over here is a 200K.
- 1:44:20And what that means is let's just
- 1:44:22hypothetically say you're at your
- 1:44:23199,000th
- 1:44:26token, okay?
- 1:44:28That means
- 1:44:29that on a equivalent query that you
- 1:44:33might have previously scored 100% at,
- 1:44:37at, I don't know, 5 or 10K tokens,
- 1:44:39at 199,000 tokens, you might only score
- 1:44:4240% at.
- 1:44:43Now, these numbers I basically pulled
- 1:44:45out of my ass to be clear, but the point
- 1:44:48I'm trying to make is the longer the
- 1:44:49token count, basically the bigger the
- 1:44:52context
- 1:44:53length,
- 1:44:55the lower the performance of the model.
- 1:44:57And so, understanding context windows
- 1:44:59and then learning a little bit of
- 1:45:01context management, ways to proactively
- 1:45:02manage all of these things, some of
- 1:45:04which you have control over and some of
- 1:45:06other things which you don't, is very
- 1:45:08important.
- 1:45:09It's also important, of course, because
- 1:45:10of billing.
- 1:45:12The more tokens that you use up,
- 1:45:14obviously the more money that you spend.
- 1:45:17And so, not only is it best from a
- 1:45:18quality perspective over here to try and
- 1:45:22push to the left side of this graph as
- 1:45:23much as humanly possible, it's It's very
- 1:45:25relevant from a financial perspective
- 1:45:28over here because obviously the more
- 1:45:30tokens that you use the more money you
- 1:45:31spend. Case in point, just to make this
- 1:45:33video I've spent something around $500
- 1:45:36or so in tokens. Now that's because I'm
- 1:45:38using a particular agent's fast mode
- 1:45:41which bills me directly instead of just
- 1:45:43using a monthly plan. But the point
- 1:45:45remains, any sort of serious AI agent
- 1:45:47application will start spending and
- 1:45:49using a fair amount of your money.
- 1:45:51Okay, and just before I move on, I want
- 1:45:53to talk a tiny bit about the differences
- 1:45:56between each of these in length. If I
- 1:45:58open up an actual Claude instance here,
- 1:46:01I open up one of these and then I go to
- 1:46:03terminal
- 1:46:04which is the current best way to
- 1:46:06visualize this. And if I just maximize
- 1:46:09this panel size, let's just make this as
- 1:46:11big as humanly possible,
- 1:46:13then I go {slash} context,
- 1:46:15Claude will show us all of the things
- 1:46:18currently consuming its tokens.
- 1:46:20I'm going to zoom in here to make it
- 1:46:21really really clear what's going on.
- 1:46:25This right over here is your context
- 1:46:27usage. And as you can see, they've
- 1:46:29illustrated this as sort of a series of
- 1:46:32squares where every square is, I don't
- 1:46:34know, let's see, 1 2 3 4 5 6 7 8 9 10.
- 1:46:37Okay, every square is, I think, 2,000
- 1:46:40tokens or so.
- 1:46:41And so what we're seeing is, despite the
- 1:46:44fact that we have put no conversation
- 1:46:48tokens, so we haven't spent any tokens
- 1:46:51at all on conversation,
- 1:46:53we're still
- 1:46:55at 9,000
- 1:46:57used.
- 1:46:59You're probably wondering where the hell
- 1:47:00are these 9,000 coming from? Are they
- 1:47:01shadow billing me to try and
- 1:47:04rinse my wallet as much as humanly
- 1:47:06possible?
- 1:47:07Well, a little bit. I mean, this system
- 1:47:09prompt here, okay, which is partially
- 1:47:12composed by your agents, your Gemini or
- 1:47:14your Claude and MD and partially a few
- 1:47:15additional things, it's actually already
- 1:47:17consuming 4,900 tokens. So 2.5% of my
- 1:47:21entire token count before I even send a
- 1:47:22message is being used by in this case
- 1:47:25probably the Claude.md.
- 1:47:29But in addition you have other things
- 1:47:30like memory files which are consuming
- 1:47:312,000 tokens, okay? And the way that
- 1:47:35they do that in Claude code is they use
- 1:47:37something called a memory.md which
- 1:47:39stores your preferences and some some
- 1:47:41previous high-level things.
- 1:47:43Then next up we have skills which are
- 1:47:44consuming 1,700 tokens. What are these
- 1:47:47skills? Well, you guys remember when we
- 1:47:49made a bunch over here? If I go to the
- 1:47:51top left-hand corner where it says doc
- 1:47:52Claude skills, you know, every agenda
- 1:47:55coding platform has their own
- 1:47:56configuration for this stuff. But the
- 1:47:58way that it works in Claude code is, you
- 1:48:00know, you organize these workflows into
- 1:48:01these skills.
- 1:48:03Well, guess what? These skills aren't
- 1:48:04free. In order for Claude to be able to
- 1:48:07use these skills, okay? This multi-agent
- 1:48:10orchestrator, it needs to store all
- 1:48:12these tokens somewhere and then give it
- 1:48:14to the model and that's what's going on
- 1:48:15over here, right? Now the actual
- 1:48:17messages that we've used
- 1:48:19are only at eight tokens. And guess
- 1:48:20what? That's actually, I think it's um
- 1:48:22this word here context usage.
- 1:48:26I'm not entirely sure, but I think
- 1:48:28context usage just because of the way
- 1:48:29that it's broken down or maybe context
- 1:48:30usage plus this term here {slash}
- 1:48:32context, you know, is equal to eight
- 1:48:34tokens.
- 1:48:35But know that, you know, that's that's
- 1:48:37it. That's it for our whole token count.
- 1:48:39So the other 158,000 of our 200,000
- 1:48:41limit is currently free.
- 1:48:43And so I mean this is a quick and easy
- 1:48:45way obviously to visualize it inside of
- 1:48:46Claude code, but um other platforms have
- 1:48:48their own visualization mechanisms. Now
- 1:48:50next up are these MCP tools. A way to
- 1:48:52look at these MCP {underscore}
- 1:48:54{underscore} Chrome dev tools
- 1:48:56{underscore} {underscore} click. What is
- 1:48:58that? Well, this is the tool that allows
- 1:49:01Chrome to click on parts of the page.
- 1:49:05Remember earlier when we were building
- 1:49:07that little N8N flow?
- 1:49:09Well, we were doing it by clicking on
- 1:49:10various parts of the page.
- 1:49:12How about drag, right? We can drag
- 1:49:15things, get console message. These are
- 1:49:17all basically buttons in some colossal
- 1:49:21spaceship. Basically, we are in the
- 1:49:24cockpit with Claude code and we're
- 1:49:26telling it to do stuff for us. We don't
- 1:49:28know what these buttons are. It does
- 1:49:31because it's, you know, the ship
- 1:49:32technician or the navigator or whatever.
- 1:49:34And so it's clicking these buttons left,
- 1:49:36right, and center for us to do various
- 1:49:37things. That's how you can conceptualize
- 1:49:39all these MCP tools and all these
- 1:49:40skills. And what I really like about
- 1:49:41this is it breaks everything down. So
- 1:49:42here are memory files, okay, which I
- 1:49:44talked about the memory.md, the
- 1:49:46claude.md. You can see this is being
- 1:49:48contributed to in a variety of ways. We
- 1:49:50have a global claude.md, which is sort
- 1:49:52of like a very high-level one with some
- 1:49:53sparse instructions. We have a local
- 1:49:55claude.md. We actually have the memory
- 1:49:57down here and then we have all the
- 1:49:58skills. You know, what was really
- 1:49:59telling about that is despite the fact
- 1:50:01that in this diagram conversation
- 1:50:02history is like the biggest chunk of it
- 1:50:04all. Notice that in reality,
- 1:50:06conversation history for us, at least at
- 1:50:08the beginning, was nothing. You know, it
- 1:50:10was actually a tremendous number of
- 1:50:12tokens, about 10% of our entire context
- 1:50:14window, dedicated to just this little
- 1:50:16chunk.
- 1:50:17And that is a problem because if you're
- 1:50:20not careful, your claude.md with all
- 1:50:22that's rules are going to get really,
- 1:50:23really long. Your memory.md with all
- 1:50:26your preferences is going to get huge.
- 1:50:27Same thing with all the skills and tools
- 1:50:29and stuff like that. And your
- 1:50:29conversation history, when you actually
- 1:50:31do get to conversating with the model,
- 1:50:33will be very, very, very small.
- 1:50:35You know, instead of starting, I don't
- 1:50:37know, somewhere in this region here,
- 1:50:40you're actually, because the context
- 1:50:42window of all of your BS is so big, you
- 1:50:44might actually like start in the
- 1:50:45effective area over here.
- 1:50:47And obviously this is not what you want
- 1:50:49to begin an agent conversation at
- 1:50:51because if you start here, then it's
- 1:50:52obviously only downhill from there. Now
- 1:50:54this takes me to the logical question
- 1:50:56of, you know, hey Nick, what happens
- 1:50:58when you run out of context? Because
- 1:51:00obviously that's going to happen.
- 1:51:02Well, when there are 50,000 tokens,
- 1:51:05let's say, out of 200,000, okay? No
- 1:51:07problem. You are having the full
- 1:51:09conversation history with the model and
- 1:51:12the model basically gets literally every
- 1:51:14message starting from message number
- 1:51:15one, two, three, four, all the way down
- 1:51:18to I don't know, message number 25. And
- 1:51:20then what it does is it takes all of
- 1:51:22this context and then feeds it into its
- 1:51:24big neural network to generate message
- 1:51:26number 26, right?
- 1:51:28However, when we get to a certain
- 1:51:29length, right? I don't know, let's say
- 1:51:31message 50. Obviously, there's no
- 1:51:33there's no tokens left anymore in the
- 1:51:34context window. Maybe we're like 199
- 1:51:36Well, not actually, but probably be like
- 1:51:37155k
- 1:51:39out of 200k.
- 1:51:40Well, what happens is all models now
- 1:51:42have some sort of what's called like
- 1:51:44auto compact limit.
- 1:51:46I mean, basically all of them have
- 1:51:47adopted this convention, where when the
- 1:51:50number of tokens that you're using,
- 1:51:52let's say the
- 1:51:54limit is right over here.
- 1:51:57When the number of tokens that you're
- 1:51:58using gets to this point, okay?
- 1:52:01You know, it fills up,
- 1:52:03then fills up,
- 1:52:04fills up,
- 1:52:06fills up, and fills up.
- 1:52:08What happens is this triggers a
- 1:52:10mechanism called compaction,
- 1:52:12where we take all of the information
- 1:52:15here,
- 1:52:18which is I don't know,
- 1:52:19maybe like 80% or so of the whole
- 1:52:21context, and then we compress it. I want
- 1:52:24you to imagine right now there's like a
- 1:52:25a big hydraulic press type thing over
- 1:52:28here,
- 1:52:29and it's pushing all of this context
- 1:52:31down, it's squishing it. Basically,
- 1:52:33what's going to occur
- 1:52:34is we're going to erase the vast
- 1:52:36majority of this,
- 1:52:38and then, you know, instead of consuming
- 1:52:4080%, we're going to cram all that
- 1:52:42information into maybe, I don't know, 30
- 1:52:44or 40% or so.
- 1:52:47And so that is called compaction,
- 1:52:50and it occurs on a relatively regular
- 1:52:52basis across the model ecosystem.
- 1:52:55The issue with the compaction, or you
- 1:52:57know, compression, or whatever the heck
- 1:52:58you want to call it, it's the same idea,
- 1:53:00is during this summarization and like
- 1:53:03densification process, we are going to
- 1:53:05drop outputs from tools. We are going to
- 1:53:08remove some information and context that
- 1:53:11might actually be useful to you now at,
- 1:53:13you know, message 50 from message four,
- 1:53:15which might actually like eliminate a
- 1:53:16mistake. And so, because of that, you
- 1:53:18know, you are going to lose some of the
- 1:53:19quality.
- 1:53:20Um the benefit is obviously you will
- 1:53:22significantly improve the information
- 1:53:24density. And what do I mean by
- 1:53:25information density? I mean literally
- 1:53:26like if it's like um hello,
- 1:53:30how are
- 1:53:32you doing?
- 1:53:35I mean, let's say somewhere in your
- 1:53:36context you have the term you the
- 1:53:38sentence hello, how are you doing? Well,
- 1:53:40hello is actually two tokens. How is
- 1:53:42one, are is one, you is one, doing is
- 1:53:46three along the question mark. So, in
- 1:53:48total, if you count all the stuff up,
- 1:53:50depending on the uh thing you're using
- 1:53:51to do the embedding, that's uh eight,
- 1:53:53right?
- 1:53:54Well, you know, compaction is literally
- 1:53:57going to take this sentence and then
- 1:53:58it's going to compress it. So, it's
- 1:53:59going to say hi,
- 1:54:01how are you?
- 1:54:04So, now instead of that's one token
- 1:54:06here, one token here, one token here,
- 1:54:08one token here equals four.
- 1:54:10It will literally get all of the context
- 1:54:12that it can and then try and squish it
- 1:54:14so that the same meaning is available in
- 1:54:16fewer tokens and fewer words wherever
- 1:54:18possible. And then it'll just run this
- 1:54:20naively across your entire um context.
- 1:54:23And so, this is going to occur every
- 1:54:24single time. Obviously, it's something
- 1:54:25that we want to avoid occurring if we
- 1:54:27have sensitive and important data, but
- 1:54:29you know, it does allow us to continue
- 1:54:31conversating with models. Before we had
- 1:54:33context compression in some form of auto
- 1:54:34comp action, um basically we would just
- 1:54:37run out of tokens and we'd have to
- 1:54:38restart a totally new session. So, this
- 1:54:40just does sort of that intermediate step
- 1:54:41that most people were doing before where
- 1:54:42they would like like take all of that,
- 1:54:44you know, try and summarize it in some
- 1:54:47other model and then paste it back into
- 1:54:48uh into into another one. Now, that
- 1:54:50takes us to what most model
- 1:54:52practitioners are using nowadays, which
- 1:54:54is a variant of something that uh people
- 1:54:57call the iceberg technique.
- 1:54:59Now, in case you guys have never seen an
- 1:55:01iceberg before, I'm Canadian, so we have
- 1:55:03them everywhere, including literally in
- 1:55:05the river across the street from my
- 1:55:07house.
- 1:55:08Not uh actual icebergs, but little ice
- 1:55:10floats.
- 1:55:11The way that they work is basically you
- 1:55:13have um a section that's visible above,
- 1:55:16which is usually quite massive and quite
- 1:55:17intimidating. And then you're like, "Oh
- 1:55:19my god, that's a really big iceberg."
- 1:55:21And then what you don't realize is
- 1:55:22underneath the iceberg is actually like
- 1:55:24two or three times as big.
- 1:55:26And so above, the stuff that's like
- 1:55:28immediately visible,
- 1:55:30or in our terms for the model,
- 1:55:33accessible immediately,
- 1:55:36is actually only a very small percentage
- 1:55:38of the total iceberg.
- 1:55:42And so in the context of our model, what
- 1:55:44we store above, which is immediately
- 1:55:46accessible, it's the sort of the stuff
- 1:55:47that's like visible to the plain eye, is
- 1:55:50we'll store our memory,
- 1:55:52we'll store our Claude or our agents or
- 1:55:54our Gemini.md,
- 1:55:56we'll store our local memory as well.
- 1:55:59There's different types, there's global,
- 1:56:00local. We'll store all of the current
- 1:56:02task contexts, everything that our tools
- 1:56:05are doing, then maybe any active file
- 1:56:07contents, okay? And so all of this stuff
- 1:56:09here is like basically always accessible
- 1:56:11to us. It's literally just in our
- 1:56:12prompt.
- 1:56:13But then, what people are doing to
- 1:56:15reduce the total number of tokens that
- 1:56:16they require, is they're abstracting
- 1:56:18away everything else. And then they're
- 1:56:19just making it accessible to the model
- 1:56:21if it needs it.
- 1:56:22What I mean by that is instead of
- 1:56:24putting all of the files in your code
- 1:56:26base in a prompt, what it does is it
- 1:56:28gives you a tool called read. What read
- 1:56:30can do is read can at any time read a
- 1:56:32file, okay? But instead of putting the
- 1:56:35entire file in there, for now all it
- 1:56:36does is just puts the titles.
- 1:56:38So if you say, "Hey, I want you to grab
- 1:56:40the um information on the iceberg
- 1:56:42technique." And then in your workspace
- 1:56:44you have a file called iceberg
- 1:56:45technique.md,
- 1:56:46it'll know it doesn't actually have to
- 1:56:47read all of the files in your workspace,
- 1:56:49it only has to read iceberg.md, right?
- 1:56:52Same thing with the full code base. You
- 1:56:53have tools like grep and glob. And these
- 1:56:56tools are sort of analogous. Instead of
- 1:56:58reading a whole file, what this does is
- 1:57:00allows you to hone in on a specific
- 1:57:01segment of text. You know, if this is my
- 1:57:03entire code file, okay, and
- 1:57:05hypothetically, let's just say it's
- 1:57:07really big and, you know, there's
- 1:57:09there's a lot of stuff. But the only
- 1:57:11thing that I actually care about is this
- 1:57:13little segment over here,
- 1:57:16then why would I load all of the, I
- 1:57:19don't know, 10K tokens?
- 1:57:21I don't need to, okay? Realistically,
- 1:57:24what I can do as a smart model is I
- 1:57:26could use grep and glob, these tools, to
- 1:57:29maybe hone in only on this segment over
- 1:57:31here, which is, I don't know, let's just
- 1:57:33say 2K tokens.
- 1:57:35And because it contains some of the text
- 1:57:37before and some of the text after, you
- 1:57:39know, this still usually gives us enough
- 1:57:40context to
- 1:57:42um tell the model what it needs in order
- 1:57:44to to finish its function. Then you also
- 1:57:45have web data via web fetch. Web data's
- 1:57:48pretty cool. You can kind of think of it
- 1:57:50as the same thing. Obviously, it doesn't
- 1:57:51have access to the whole internet, but
- 1:57:52it can make search queries, right? And
- 1:57:54so, because it's able to use some
- 1:57:56general reasoning, when you say, "Hey,
- 1:57:57what's the iceberg technique?" first
- 1:57:59it's going to start by looking for file
- 1:58:01contents called iceberg technique. If it
- 1:58:03can't find any, maybe it'll quickly type
- 1:58:05F iceberg to look through the code base.
- 1:58:07If it can't find that, you know, maybe
- 1:58:09it'll look through some other things
- 1:58:10like uh memory files in the skills
- 1:58:12library. But if it can't find that,
- 1:58:14it'll say, "Okay, cool. So, we don't
- 1:58:15have this in our in our context window
- 1:58:17right now, and we don't even have access
- 1:58:19to it within our space, but it's
- 1:58:20probably somewhere on the internet, so
- 1:58:21I'm just going to Google iceberg
- 1:58:22technique." And then what it'll do is it
- 1:58:25won't even take the entire thing. It'll
- 1:58:26just grab top-level links, and then
- 1:58:28it'll look at the URL, you know, uh and
- 1:58:31the URL is about, guess what, icebergs.
- 1:58:33I'm going off the map here, but
- 1:58:35hopefully you guys understand what I
- 1:58:36mean. Then if there's three URLs, one of
- 1:58:38them's about icebergs, then instead of
- 1:58:39reading all of them, it's only going to
- 1:58:40read this one. And so, it's sort of like
- 1:58:42a successive narrowing of the lens until
- 1:58:45eventually it gets to, you know, what
- 1:58:47you want. And it doesn't load the the
- 1:58:50of all of the context, but it just has
- 1:58:53the opportunity select all of this. You
- 1:58:55know, it starts here, it goes here, goes
- 1:58:57here, goes here, goes here, goes here,
- 1:58:58and then goes here, and then finally
- 1:58:59it's at its goal.
- 1:59:01You can do the same thing in a variety
- 1:59:02of other ways. You can use bash, get
- 1:59:03history, and so on and so forth, but
- 1:59:06um essentially what you want to do is
- 1:59:08instead of storing all of the
- 1:59:09information like the full code base, the
- 1:59:11file contents, the web data, all the get
- 1:59:13history, all the skills, and everything
- 1:59:14like that. What you do is you just store
- 1:59:16the ability to access it on demand.
- 1:59:19And then inside the context, okay, the
- 1:59:22tiny chunk that you do, I mean in this
- 1:59:23diagram it says 10% 90%, you know, I
- 1:59:25think in reality it's probably closer to
- 1:59:272080 or maybe 3070.
- 1:59:30Um here's where you store stuff that's
- 1:59:31just like
- 1:59:32it it needs to be the same all the time.
- 1:59:34It always needs to be present. There
- 1:59:36always needs to be some sort of patterns
- 1:59:38that are learned, some sort of active
- 1:59:39file contexts or
- 1:59:41contexts or current task context. You
- 1:59:43can think of this as the difference
- 1:59:44between naive versus strategic context
- 1:59:48loading. Now way back in the day, and
- 1:59:50when I say way back in the day, I mean
- 1:59:52like 2023. Good God, I'm getting old.
- 1:59:55You know, when you were working with an
- 1:59:56agent, you would dump in the whole code
- 1:59:57base. You would honestly just copy and
- 2:00:00paste everything and just hope to God
- 2:00:01that it that it knew what it was doing.
- 2:00:03Obviously, because this was
- 2:00:05extraordinarily infeasible, there were
- 2:00:08so many tokens in the context.
- 2:00:11Tons of it was lost, and you routinely
- 2:00:13ran out of context limits.
- 2:00:15Well, nowadays what we've done is we
- 2:00:16basically built in a whole tool stack
- 2:00:18where instead of all of the file, you
- 2:00:20just read selectively and only the
- 2:00:22relevant functions.
- 2:00:23You know, you have a cloud.net which is
- 2:00:25basically a compression a compression
- 2:00:27function which is stores your
- 2:00:28preferences.
- 2:00:29Skills instead of all being read, you
- 2:00:31know, we only read a specific segment of
- 2:00:33them. This is technically called a YAML
- 2:00:35front matter, which is just a tiny
- 2:00:37little section at the beginning of the
- 2:00:38skill. You can actually see this if we
- 2:00:40go back to the skill.net and I make this
- 2:00:42visible, right? Only this section up
- 2:00:44here is actually and know, can actually
- 2:00:46see this up here on the YAML front
- 2:00:48matter is this segment. Only this
- 2:00:51segment is actually loaded into context
- 2:00:53until you ask for, you know, more
- 2:00:55information about create proposal. And
- 2:00:57that's just because this little space
- 2:00:58invader sees you it has the ability to
- 2:01:00call a create proposal skill, but it
- 2:01:02doesn't need to know all the rest of
- 2:01:04this because what's realistically going
- 2:01:05to happen well, 90% of the time you
- 2:01:06won't even ask. There's so many other
- 2:01:07skills that it'll probably be using.
- 2:01:09We're also now doing things like uh
- 2:01:10summarizing tool results. So, instead of
- 2:01:13storing the entire thing in your
- 2:01:15context, you know, we just store like a
- 2:01:17very short summary of basically the
- 2:01:18input output. So, way back in the day I
- 2:01:20used to think that all of an agent was
- 2:01:23really just its intelligence. It was
- 2:01:25just the core model, right? Which in
- 2:01:27this case would be Opus 4.6.
- 2:01:30But what I've come to quickly realize is
- 2:01:32although models themselves are quite
- 2:01:34intelligent, it's really the
- 2:01:36architecture that we've wrapped around
- 2:01:37it. I want you to pretend this little
- 2:01:39space invader is kind of now in I don't
- 2:01:42know a
- 2:01:43house of some kind. The house has a
- 2:01:45little chimney with a fireplace. It has
- 2:01:47a little place where it could I don't
- 2:01:48know, roast a nice turkey. You know, it
- 2:01:51has a nice bed that it can go to sleep
- 2:01:52in every night. You know, this agent by
- 2:01:55itself probably wouldn't last super long
- 2:01:57out there on the savanna, but because
- 2:01:58we've built all this infrastructure
- 2:02:00around it, because we built roads for
- 2:02:01it, we built ways for it to communicate
- 2:02:03and stuff like that, you know, it's it's
- 2:02:05capable of actually doing a lot of very
- 2:02:06economically valuable work for us.
- 2:02:09And so, just like human beings
- 2:02:11way back in the day had to conceptualize
- 2:02:13the idea of
- 2:02:14I don't know, like a like a spear or
- 2:02:16something to hunt um um
- 2:02:19saber-tooth tigers on the plains, you
- 2:02:22know, so too do these agents use tools
- 2:02:26effectively to solve problems in their
- 2:02:28environment and ultimately get us the
- 2:02:30users what we want. Now, obviously
- 2:02:32there's context window management like
- 2:02:35we just talked about. And that's for
- 2:02:37optimizing the usage of a specific
- 2:02:40model, okay, in the choice of a model.
- 2:02:43but there's also the ability to choose
- 2:02:44different models for different purposes.
- 2:02:47Now, throughout most of the course so
- 2:02:49far, what I've done is I've just used
- 2:02:51mostly naive like Opus 4.6 agents in
- 2:02:54order to spawn other Opus 4.6 sub
- 2:02:56agents.
- 2:02:57And that was mostly for capability sake
- 2:02:59because at least in my case cost is not
- 2:03:01a a concern and I really do want to eke
- 2:03:03out the marginal quality um benefits
- 2:03:05wherever possible.
- 2:03:06But there are a lot of cases,
- 2:03:08specifically like enterprise and and big
- 2:03:10infrastructure ones, where, you know,
- 2:03:11people are actually comfortable making a
- 2:03:13minor trade-off for quality. I'm going
- 2:03:16to draw another one of my famous graphs.
- 2:03:18Um people are comfortable making, you
- 2:03:20know, a a a trade-off in terms of, you
- 2:03:22know, cost
- 2:03:24and then quality. Now, in a lot of
- 2:03:26disciplines out there, biology, physics,
- 2:03:28chemistry, and stuff, there's this this
- 2:03:30idea of like this inverted U curve. And
- 2:03:33you can call it whatever the heck you
- 2:03:34want. Um I think the actual name is the
- 2:03:37Yerkes-Dodson curve.
- 2:03:40And what this is is this is basically
- 2:03:41like the the optimal point
- 2:03:45that combines two different factors. And
- 2:03:48so in our case, if we're optimizing for
- 2:03:49both cost and quality, okay,
- 2:03:51simultaneously, not just cost, not just
- 2:03:53quality, you can imagine that the
- 2:03:55optimal place to choose on this graph is
- 2:03:57probably going to be something like over
- 2:03:59here. It's going to be It's going to be
- 2:04:00somewhere around here.
- 2:04:01If you wanted to minimize cost exactly,
- 2:04:03we'd probably go like way over here. But
- 2:04:05obviously we care about quality as well,
- 2:04:06so we're going to push up a little bit,
- 2:04:08right? We don't want this point because
- 2:04:11one, it costs a lot, and then boom, two,
- 2:04:13the quality's pretty low, because
- 2:04:15despite the fact that quality is really,
- 2:04:16really high, um you know, cost is like
- 2:04:18many, many times higher than it was over
- 2:04:20here.
- 2:04:21And so this is sort of like our our
- 2:04:23optimal point. And, you know, a lot of
- 2:04:25large enterprises, since we're dealing
- 2:04:26with hundreds of millions of dollars
- 2:04:28here, um are actually comfortable making
- 2:04:30a little trade-off where, you know, if
- 2:04:31this quality is 85% and then this one
- 2:04:34here is, I don't know, 95%, they're okay
- 2:04:36taking like a like a 10% hit here
- 2:04:40if it means that they also reduce their
- 2:04:42costs by, I don't know, 40% or something
- 2:04:44like that.
- 2:04:46And this is really where all this stuff
- 2:04:47comes in, okay? So, I don't mean to talk
- 2:04:48your ear off here. It's not super
- 2:04:49important, but uh basically uh what a
- 2:04:52lot of people have taken to doing now is
- 2:04:53doing a 60-30-10 rule
- 2:04:55where they'll use some top-level agent
- 2:04:58router, which is sort of like the
- 2:05:00orchestrator in our multi-agent uh
- 2:05:02Chrome window example. And what that
- 2:05:05agent router does is it calls dumber
- 2:05:06models
- 2:05:08and then it assigns different strengths
- 2:05:10to tasks so that, you know, if you're
- 2:05:13giving it a really simple task and you
- 2:05:15say, "Hey, you know, I just want you to
- 2:05:16classify this into one of three
- 2:05:18categories."
- 2:05:19And it's really dumb. It's like, you
- 2:05:21know, red, blue, or green,
- 2:05:23angry, serene, or or healthy, or
- 2:05:26whatever the heck. Um you don't have to
- 2:05:28use like Opus, which is like space-age
- 2:05:30intelligence and costs you a ton more in
- 2:05:32token cost in order to get that done.
- 2:05:35Instead, you know, you can get all the
- 2:05:36way down to like a Haiku or maybe a
- 2:05:38Gemini Flash model or something like
- 2:05:39that. Likewise, if you have some other
- 2:05:41task here, and maybe that task requires
- 2:05:43a lot of um I don't know, research or
- 2:05:45something,
- 2:05:46and you say, "Hey, I want you to go and
- 2:05:47compile 200 million tokens worth of
- 2:05:50stuff and then um give it to me in a big
- 2:05:52report." Well, you know, you probably
- 2:05:53don't want the dumbest model to do that
- 2:05:55for you. But you also don't need the
- 2:05:55most expensive model. So, maybe you'll
- 2:05:57use something like a Saw model or like a
- 2:05:59a lower-level GPT model, which might
- 2:06:00cost two or three million dollars uh two
- 2:06:02or three dollars per million tokens
- 2:06:04instead. And so, you know, this
- 2:06:05allocation where you have 60-30-10, if
- 2:06:08you think about it like sort of a pie
- 2:06:09chart, um what you do is you designate,
- 2:06:12you know, the vast majority of your
- 2:06:15token usage
- 2:06:16to stuff that is in that first category,
- 2:06:18which is kind of, you know, dumber.
- 2:06:21And then what you do is you do the other
- 2:06:2330% or so
- 2:06:25in sort of that mid-tier. And then your
- 2:06:27really, really, really smart models, you
- 2:06:29know, they do the highest-level tasks.
- 2:06:32And basically, what will happen for the
- 2:06:33most part is this would be, you know,
- 2:06:34your Opus 4.6 or your Gemini
- 2:06:393.1 or your GPT 5.4. It'll be
- 2:06:43responsible for routing decisions and
- 2:06:44obviously you want the smartest model
- 2:06:46possible for that. But all of the heavy
- 2:06:47lifting, all the context and stuff like
- 2:06:49that is um through sub-agents that I
- 2:06:51spawned, either Haiku,
- 2:06:53Sonnet,
- 2:06:54or uh you know, I don't know if if if
- 2:06:56you wanted to do a really smart call,
- 2:06:57then you'd obviously spawn an Opus
- 2:06:59sub-agent like me as well.
- 2:07:01And if you do all this, you could
- 2:07:02significantly reduce the cost. I mean,
- 2:07:03like just just think about it
- 2:07:04mathematically. If previously you were
- 2:07:06doing 100 million tokens times, you
- 2:07:09know, $5 per 1 million tokens,
- 2:07:12>> [gasps]
- 2:07:13>> what's the cost there?
- 2:07:14Well, that's obviously going to cost
- 2:07:15$500. And so that's like your Opus only,
- 2:07:18right? But if you did 10 million times
- 2:07:21$5 plus
- 2:07:2430 million times $3
- 2:07:28plus 60 million times
- 2:07:31$1, what's the total cost going to be
- 2:07:33now? Well, it's going to be 60
- 2:07:36plus 90
- 2:07:39plus 50
- 2:07:41or in total, 200.
- 2:07:44And so, you know, 200 expressed as a
- 2:07:45fraction of 500 is 40%
- 2:07:49of our total cost. And we will have just
- 2:07:52saved, you know, 60% with probably
- 2:07:54minimal impacts on quality because the
- 2:07:57things that we're now spawning, you
- 2:07:59know, dumber agents um to do are things
- 2:08:01that, to be honest, the quality was
- 2:08:03already okay a few generations ago back
- 2:08:05with the Haikus and the Sonnets. So just
- 2:08:07to give you an example from something
- 2:08:08that I do pretty often, okay, which is
- 2:08:11going to be some form of lead scraping,
- 2:08:13what you can do is you can actually
- 2:08:14traverse a very large portion of the
- 2:08:17internet using a relatively dumb model,
- 2:08:19these Haiku models. What these do is
- 2:08:22these um scrape a vast majority a vast
- 2:08:25amount of internet data, okay? All of
- 2:08:27the code of uh I don't know, let's say
- 2:08:2910,000 websites or something. And then
- 2:08:32in doing so, they just like that little
- 2:08:33magnifying glass, use some sort of grep
- 2:08:35or extraction prompt to look for things
- 2:08:38that are formatted like email addresses.
- 2:08:40So, if you have like something and then
- 2:08:42it's an at and then it's a you know, the
- 2:08:44term gmail.com,
- 2:08:46odds are this is a real email address,
- 2:08:47right? So, then you store that to a
- 2:08:48database. And because this is just such
- 2:08:50like a mass data application, use HiQ.
- 2:08:52It drives the cost really, really low.
- 2:08:55Well, then maybe um the actual
- 2:08:56enrichment point, you know, takes
- 2:08:58significantly more intelligence. And so,
- 2:09:00maybe here we'll use Sonnet and it'll
- 2:09:01cost us uh $0.008
- 2:09:04per lead, $0.008 per lead. The actual
- 2:09:07outreach part is mostly templated, so we
- 2:09:10use Sonnet for that as well. And then
- 2:09:11maybe at the end we just have a quality
- 2:09:13review step to make sure things aren't
- 2:09:14absolutely nuts. Well, when you do it
- 2:09:16this way, um you know, the math ends up
- 2:09:18being uh $0.008 + $0.005, so that's
- 2:09:21$0.013 + $0.001, $0.014 + $0.015 is
- 2:09:26$0.029.
- 2:09:28And then if you were to go 100% Opus,
- 2:09:30then it would be I don't know, uh uh
- 2:09:32about 12 cents or so per lead. And so,
- 2:09:34on a list of let's just say a thousand,
- 2:09:36which is approximately how much I'm
- 2:09:38sending a day right now. I'm much
- 2:09:39farther down than uh my maximum, but if
- 2:09:42we multiply all these together, we move
- 2:09:43that one, two, three decimal points over
- 2:09:45to the left, then that um ultimately
- 2:09:48would be $15 a day.
- 2:09:51Or, you know, $450 a month. Well,
- 2:09:54instead I'm doing that literally one
- 2:09:57quarter of this. Or something like, you
- 2:09:59know, $120
- 2:10:01a month instead. Obviously, I'd much
- 2:10:03rather the latter. And if my quality is
- 2:10:05only going down a few percentage points
- 2:10:06because of that Yerkes-Dodson curve, you
- 2:10:09know, I'm okay being over here instead
- 2:10:12of over here because this gap to me is
- 2:10:14fine on tasks that aren't super high
- 2:10:16high quality, then uh this is a very,
- 2:10:19very efficient stack. And you know, the
- 2:10:21bigger and bigger my company gets,
- 2:10:23whoever I'm working with, the more the
- 2:10:25cost per lead is going to be important
- 2:10:28versus the actual quality. I've included
- 2:10:30just a little LLM API pricing cheat
- 2:10:33sheet. I don't expect this to be super
- 2:10:35relevant or useful to you guys. There
- 2:10:37are a lot more models for OpenAI
- 2:10:40and Google, but I am actually using um
- 2:10:42not this model series anymore, but this
- 2:10:44one here for some queries. I'm also
- 2:10:46using the flash model series for some
- 2:10:48queries as well.
- 2:10:49And then what's really cool is a lot of
- 2:10:51them offer uh what's called a batch API
- 2:10:53now, where you can submit a bulk number
- 2:10:56of requests over um simultaneously. And
- 2:10:59then if you're comfortable waiting like
- 2:11:01a day or so, what the companies do is
- 2:11:03they batch it and then they serve your
- 2:11:06requests during periods in which they
- 2:11:08have very low inference or low
- 2:11:09competition. So, maybe in the middle of
- 2:11:11the night or something like that. And in
- 2:11:13doing so, they actually get to load
- 2:11:14balance. Like if you think about it like
- 2:11:16if this is like a day
- 2:11:18and then this is their load like on
- 2:11:19their servers and their neural networks
- 2:11:21and stuff, you know, it'll probably peak
- 2:11:22somewhere around like noon and there's
- 2:11:24probably like a couple things and then
- 2:11:25it's like low during the during the day,
- 2:11:27right? I don't know. This is like 4:00
- 2:11:28a.m. What they'll do
- 2:11:31is they'll actually take all of your
- 2:11:32queries, batch them, and then they'll
- 2:11:33just like run them over here when
- 2:11:35there's very little competition.
- 2:11:37And then later when, you know, the
- 2:11:38things go up and stuff like that again,
- 2:11:40um that's okay. And in doing so, what
- 2:11:42they want to do is they want to shift
- 2:11:43some of the the really top end of all of
- 2:11:45these users
- 2:11:47um to the low end to basically fill this
- 2:11:49so that they have a lot more like
- 2:11:50dependable load instead of these like uh
- 2:11:53jagged peaks and whatnot.
- 2:11:55>> [sighs and gasps]
- 2:11:56>> But anyway, don't worry too much about
- 2:11:57that. I just wanted to cover um some LLM
- 2:11:59pricing principles as well so that you
- 2:12:00guys know not only how to manage your
- 2:12:02context better, but also how to um save
- 2:12:06especially when you get into more
- 2:12:07sophisticated multi-agent setups like
- 2:12:09I've been showing you. And that's it.
- 2:12:11Thank you guys very much for watching
- 2:12:12this video end to end. If you guys have
- 2:12:14made it all the way to this point in the
- 2:12:15course, you're part of like the 2 or 3%
- 2:12:18that actually do. Um I'd really
- 2:12:19appreciate a big solid if you could do
- 2:12:21me a favor and subscribe to the channel.
- 2:12:23Something like 70% of you aren't, which
- 2:12:25significantly hurts my reach. And uh
- 2:12:27despite me hating asking for it, it does
- 2:12:29help the channel grow. So, if I've given
- 2:12:30you guys any value whatsoever, please do
- 2:12:32that. You can also send me over a
- 2:12:34comment down below asking any question
- 2:12:36about any point in the video. Um I'm
- 2:12:39much more engaged than the average
- 2:12:40YouTuber, so the probability that I will
- 2:12:41reply is pretty pretty high up there, I
- 2:12:43would say, statistically.
- 2:12:45Um if you guys have any, you know,
- 2:12:47suggestions for future videos or future
- 2:12:48courses, please drop them down below as
- 2:12:50well. And above all else, keep learning
- 2:12:52and growing with AI agents. This is by
- 2:12:55far the biggest and most impactful of
- 2:12:58economic changes that I think any of us
- 2:13:00will see in our lifetime. It's a blessed
- 2:13:02time to be alive in general. You might
- 2:13:03as well not waste it. Make the most out
- 2:13:05of it. All right. Um thank you very
- 2:13:07much. Feel free to use the chapter
- 2:13:08headings to revisit any section in the
- 2:13:10course. And uh looking forward to seeing
- 2:13:11all y'all in the next one. See you
- 2:13:13later.
About this transcript
This page contains the full transcript of AI Agents Full Course 2026: Master Agentic AI (2 Hours) by Nick Saraev, generated from the public captions YouTube serves with the video. The transcript has 28,280 words across 4,226 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.