YouTube2Text

AI Agents Full Course 2026: Master Agentic AI (2 Hours) — Transcript

by Nick Saraev · 28,280 words · 4,226 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hey, this is the definitive course on AI
  2. 0:02agents. I currently teach over 2,000
  3. 0:04people how to use AI agents in both
  4. 0:05their personal and business lives and
  5. 0:07run a business that does over $4 million
  6. 0:09a year using AI agents. So, you don't
  7. 0:11need any programming or pre-existing
  8. 0:13computer experience in order to make
  9. 0:14this course work for you. I myself don't
  10. 0:16have a formal computer science degree.
  11. 0:18I've learned everything that I know
  12. 0:19watching free resources like you doing
  13. 0:21now. This is also a general AI agents
  14. 0:23course, so you don't need to know any
  15. 0:25specific platform. This isn't just on
  16. 0:27Codex or Claude Code or Anti-Gravity,
  17. 0:29but rather on all of them. So, wherever
  18. 0:31you guys are starting, you'll end up at
  19. 0:33the same place. No fluff, here's what
  20. 0:34you're going to learn in this course.
  21. 0:35First, I'll show you guys a demo where
  22. 0:37I'm controlling five AI agents, each
  23. 0:39with their own Chrome browsers as they
  24. 0:41interact with the web and perform
  25. 0:42economically valuable activities for me.
  26. 0:44I wanted to front-load this course with
  27. 0:46a demo so you guys could see what we're
  28. 0:47working up to. And just a few months
  29. 0:48ago, what I'm doing here would have been
  30. 0:50considered absurd. Then, I'm going to
  31. 0:51cover the core AI agent workflow loop,
  32. 0:53which works independent of which
  33. 0:54platform you're using. After that, I'm
  34. 0:56actually going to talk about and then
  35. 0:57sign up to the three major AI agent
  36. 0:59platforms right now. So, I'll sign up to
  37. 1:00Codex, to Anti-Gravity, and then Claude
  38. 1:02Code. And then after I'll cover what
  39. 1:04each platform is at the moment the best
  40. 1:06or the worst at. Then, we're going to
  41. 1:07dive into foundational AI agent
  42. 1:09prompting techniques. So, self-modifying
  43. 1:11agent instructions where the agent will
  44. 1:13rewrite its own rules to minimize the
  45. 1:14number of errors made. Multi-agent MCP
  46. 1:17orchestration, which is where we'll
  47. 1:18register Codex, Gemini, and Claude as
  48. 1:20MCP servers so you can manage multiple
  49. 1:22agents within a single conversation
  50. 1:24thread. Video-to-action pipelines where
  51. 1:26we'll teach agents to learn from YouTube
  52. 1:28videos instead of plain text alone.
  53. 1:30Stochastic multi-agent consensus where
  54. 1:31we'll spawn agents with the same prompt
  55. 1:34and then use their statistical spread in
  56. 1:36order to ideate and improve things
  57. 1:37about. Agent chat rooms where you'll
  58. 1:39build centralized places for agents to
  59. 1:41debate ideas, pushing them to much
  60. 1:43higher quality answers than before.
  61. 1:45Sub-agent verification loops where your
  62. 1:47agents will actually review each other's
  63. 1:48work in real time to catch things that
  64. 1:50one of them might have missed.
  65. 1:52We'll talk prompt contracts. I'll show
  66. 1:53you guys reverse prompting and a bunch
  67. 1:55of other techniques as well. And
  68. 1:57finally, we'll chat about context
  69. 1:58management and improving the agent
  70. 2:00output quality before closing out by
  71. 2:02discussing how to optimize AI agent and
  72. 2:05then token pricing. So far, I haven't
  73. 2:06seen anybody on YouTube discuss most of
  74. 2:08what I cover in this course. So for all
  75. 2:09intents and purposes, you guys consider
  76. 2:11this the sauce. Please bookmark this
  77. 2:13video, subscribe to the channel, and
  78. 2:14let's get into it. First, I want to show
  79. 2:16you how powerful these agents can be
  80. 2:18when you learn how to distribute work
  81. 2:20across multiple Chrome instances and
  82. 2:22give each sub-agent their own workspace.
  83. 2:24What I have here is a simple list of
  84. 2:27leads from, let's just say, a
  85. 2:29conference. Now, we have fields like
  86. 2:31their websites, their LinkedIn
  87. 2:33description, their first name, their
  88. 2:34last name, but one thing is missing:
  89. 2:37their email address. Now, just a year
  90. 2:39ago or so, that would have invalidated
  91. 2:42my ability to reach out to these leads.
  92. 2:44But now, because I possess their
  93. 2:46websites, I can actually spawn a bunch
  94. 2:48of Claude code agents, have them go to
  95. 2:50the websites, then have them
  96. 2:51interactively and dynamically fill out
  97. 2:53their contact forms.
  98. 2:55So what just happened as I was talking
  99. 2:57was Claude went ahead and then opened up
  100. 2:59a bunch of different Chrome browsers for
  101. 3:00me.
  102. 3:01I'm going to rearrange these to make it
  103. 3:02really easy to see. And so, this might
  104. 3:04be a little bit tough to see, but what
  105. 3:06these agents are all doing is they're
  106. 3:07independently navigating over to the
  107. 3:10contact fields of each of these
  108. 3:11websites. They're then dynamically
  109. 3:13filling out fields like the first name,
  110. 3:16the last name, the email address, and so
  111. 3:18on and so forth. And then they're
  112. 3:19putting in a little bit of outreach
  113. 3:22that's templated, but then changes
  114. 3:23depending on who they're reaching out
  115. 3:25to. These agents, through a combination
  116. 3:26of both research and then communication
  117. 3:28between each other in a shared chat
  118. 3:30room, are capable of doing things that
  119. 3:32any one agent might have taken many,
  120. 3:34many hours to do before. This is what
  121. 3:35I'm going to work up to with you guys
  122. 3:37over the course of the rest of the next
  123. 3:39couple of hours. The main strength of AI
  124. 3:41agents is really their ability to
  125. 3:43parallelize, which is to run multiple
  126. 3:46instances of each of them simultaneously
  127. 3:48while they accomplish a task. Now, right
  128. 3:51now, I would say most AI agents aren't
  129. 3:53as intelligent or as capable as a human
  130. 3:55being for any given need. But, what they
  131. 3:58are much better at us than is being
  132. 4:00fast. And so, despite the fact that
  133. 4:03their accuracy might be a little bit
  134. 4:04lower than a human, their ability to
  135. 4:06one-shot stuff is worse than ours at the
  136. 4:09moment, they can run multiple instances
  137. 4:11of themselves simultaneously and try
  138. 4:13multiple approaches over and over and
  139. 4:15over and over again in order to
  140. 4:17ultimately achieve much better results
  141. 4:19than we can. The key is you need to know
  142. 4:21a little bit about how they work under
  143. 4:23the hood. Then, you need to be able to
  144. 4:24combine them using elaborate prompt
  145. 4:26architecture like I'm going to show you
  146. 4:28in this course. So, why don't we start
  147. 4:29with one of the simplest, most
  148. 4:30foundational concepts before I actually
  149. 4:32guide you guys through signing up and
  150. 4:34setting up these different agents. And I
  151. 4:36call this the core agent loop.
  152. 4:39To make a long story short, I think most
  153. 4:41of you probably have intuition about how
  154. 4:43agents do things, but really what
  155. 4:45they're doing at the end of the day is
  156. 4:48they're going through a loop over and
  157. 4:50over and over again. And this loop is
  158. 4:52composed of three major functions.
  159. 4:55The first is the observation step. And
  160. 4:58so, here the agent is basically reading
  161. 5:00through all of its context. We're going
  162. 5:02to chat a little bit more about how to
  163. 5:04optimize and manage that later. That
  164. 5:06includes things like its files, its
  165. 5:08previous tool calls, it includes all of
  166. 5:11the system prompts, the Claude, Gemini,
  167. 5:13and agents.mds that you provide. If it
  168. 5:16does research in a previous step, it'll
  169. 5:18include the research from the internet.
  170. 5:20Uh if you're feeding in multimodal data
  171. 5:22like vision data, camera data,
  172. 5:24uh you know, audio files, and so on and
  173. 5:26so forth, it'll include all of that. And
  174. 5:28so, this agent, okay, is just in an
  175. 5:30environment and it's just always
  176. 5:32observing what's going around it, at
  177. 5:34least to start, in the observation step.
  178. 5:37From there, it'll reason. And so, this
  179. 5:39is the think step. Here, it'll consider,
  180. 5:42based off of all of this context and
  181. 5:44based off of, you know, the user's
  182. 5:45high-level goal, what do I do next? How
  183. 5:47should I plan my approach? And nowadays,
  184. 5:50most agentic coding platforms make use
  185. 5:52of like a dedicated reasoning step that
  186. 5:55you can actually click into and see,
  187. 5:57which I'll show you guys a little bit
  188. 5:58more of. And this provides a tremendous
  189. 6:00amount of interpretability,
  190. 6:01accountability, and then steerability,
  191. 6:03which is really important that I think
  192. 6:04most people sleep on.
  193. 6:06After it's thought about things and
  194. 6:07basically wrote its own mini plan, it's
  195. 6:09time to actually act, right? And so
  196. 6:12here's where it'll call tools. It'll
  197. 6:14edit the files that it decided to uh do
  198. 6:16so earlier in the plan. Or maybe it'll
  199. 6:18run a command using command line
  200. 6:20interfaces, CLIs.
  201. 6:23After the action step is done, what it
  202. 6:25does is it gets the result of the tool
  203. 6:28call, and then it feeds all of that
  204. 6:30stuff back in to the observe step. So
  205. 6:32now we're basically running through that
  206. 6:34loop again, just with a little bit more
  207. 6:36context. And so what occurs essentially
  208. 6:39is we just tend to grow bigger and
  209. 6:41bigger and bigger and bigger. If our
  210. 6:43initial context was a certain size, our
  211. 6:46you know, second loop, it's a little bit
  212. 6:48bigger. Our third loop, it's a little
  213. 6:49bit bigger. And fourth loop and so on
  214. 6:51and so forth. And what this is doing is
  215. 6:53this is basically stacking uh more and
  216. 6:55more tokens into the context that the
  217. 6:58model can then use to plan its next
  218. 6:59step.
  219. 7:01What occurs after you go through this
  220. 7:02loop, you know, usually three or four
  221. 7:04times, is eventually the model reaches a
  222. 7:06point called the definition of done.
  223. 7:11And what the definition of done is,
  224. 7:13which I think a lot of people leave out
  225. 7:15of their agent prompts, which is
  226. 7:16probably why they're always underwhelmed
  227. 7:17by what happens, is it's the series of
  228. 7:19constraints and technical specifications
  229. 7:22required for the model to conclude that
  230. 7:26it no longer needs to do this loop.
  231. 7:29Once it reaches this definition of done,
  232. 7:31okay, over and over and over and over
  233. 7:32again,
  234. 7:33it notices and then it changes routes.
  235. 7:36So now it goes to the task complete
  236. 7:38route, where it generates a quick little
  237. 7:40final response for the user. Usually
  238. 7:42involves a nicely formatted answer, as
  239. 7:44I'm sure you guys know. Hey Nick, just
  240. 7:46finished your new thumbnail app build.
  241. 7:50And before outputting it in a window,
  242. 7:52either in antigravity or Codex or maybe
  243. 7:55Claude code, in a packaged way that you
  244. 7:57guys are familiar with.
  245. 7:59And so, obviously, if you have any
  246. 8:01intuition about how AI works at this
  247. 8:04point, if you've ever communicated with
  248. 8:06ChatGPT or, you know, Claude or some
  249. 8:09other sort of desktop AI that's nestled
  250. 8:11into another application that you guys
  251. 8:12use, you'll probably know some of this
  252. 8:14stuff um just as like the foundation.
  253. 8:16But I wanted to make it really explicit
  254. 8:18at the beginning of this course because
  255. 8:20we're going to return to each of these
  256. 8:21steps over and over and over again. And
  257. 8:23it turns out that you can heavily
  258. 8:25optimize all three of these. You can
  259. 8:27optimize the hell out of the observe
  260. 8:28step. You can optimize the hell out of
  261. 8:31the think step. And understandably, you
  262. 8:33can optimize the hell out of the act
  263. 8:34step as well. That's what we're going to
  264. 8:36learn. Another point I'm going to make
  265. 8:37in this course is that AI agents aren't
  266. 8:40just the large language models
  267. 8:41themselves. You know, I think neural
  268. 8:43networks and transformers are obviously
  269. 8:46super inherently interesting because
  270. 8:48they're these massive statistical things
  271. 8:50and these beings that can that can do
  272. 8:52things. They can reason. They're very
  273. 8:54far removed from traditional computer
  274. 8:55programs just 5 or 10 years ago. So a
  275. 8:58lot of interest goes to the LLM. But I
  276. 9:00want you guys to know that the LLM
  277. 9:02really is just a very small part of what
  278. 9:04most people consider AI agents these
  279. 9:06days.
  280. 9:06The LLM is of course your reasoning
  281. 9:08engine, right? Of course it understands
  282. 9:11language and of course it makes
  283. 9:12decisions. But it's kind of like a human
  284. 9:14being from like 20,000 years ago with
  285. 9:17like a spear in its hands, right?
  286. 9:19Without all of the infrastructure around
  287. 9:21human beings, without like your your
  288. 9:23house and your fireplace and your hearth
  289. 9:25and a place to sleep at the end of the
  290. 9:26night and a a society where people farm
  291. 9:29and produce resources and you have cars
  292. 9:32that you can get in and traverse a lot
  293. 9:33of distance. Without all the tools and
  294. 9:34the architecture around the
  295. 9:36intelligence, the intelligence is
  296. 9:37actually quite limited in what it can
  297. 9:39do.
  298. 9:40And that's where the rest of these
  299. 9:41sections come into play. So, tools, much
  300. 9:44like human beings, have the ability to
  301. 9:46read files, run code, search the web,
  302. 9:49call APIs, and edit files, okay? So,
  303. 9:52too, does this AI agent. Much like human
  304. 9:54beings have the ability to set a
  305. 9:56high-level goal and keep going until
  306. 9:59that task or goal is reached, you know,
  307. 10:01so, too, can agents. And much like human
  308. 10:03beings have some sort of persistent
  309. 10:05memory where we can keep track of things
  310. 10:07that we've done and then realize that
  311. 10:09some of those things didn't work, so we
  312. 10:11got to take a slightly different tack
  313. 10:12the next time, so, too, agents have
  314. 10:14things like agents.md, claw.md,
  315. 10:16gemini.md, access to their conversation
  316. 10:19history, access to auto memory files,
  317. 10:21and skills. And so, it's not actually
  318. 10:23just the LLM, for instance, that makes
  319. 10:26an agent work. It's really all of these
  320. 10:28things multiplied by the fact that, you
  321. 10:30know, the LLM provides us like the
  322. 10:32ability to be a little bit flexible. And
  323. 10:33that's the really big different from
  324. 10:35just, you know, a chatbot and then an AI
  325. 10:37agent. A chatbot might just be the LLM,
  326. 10:40okay? But, an agent takes that that LLM
  327. 10:43and then it adds on tools, a reasoning
  328. 10:45loop, memory, and so on, and so on, and
  329. 10:47so forth. So, as a brief example, I'll
  330. 10:49use an agent coding platform called
  331. 10:51Codex. And down here, I have a simple
  332. 10:54prompt where basically, I just want this
  333. 10:56to do a bunch of research for me on
  334. 10:57creatine supplementation in men.
  335. 11:00And what I'm doing is I'm giving it a
  336. 11:01brief definition of done where I'm
  337. 11:03saying it once you've compiled 10 plus
  338. 11:05empirical sources, return a structured
  339. 11:07report. And I'm doing this cuz I want to
  340. 11:08demonstrate this loop to you. And so,
  341. 11:10there are a bunch of other things that
  342. 11:11are popping up here. We have the actual
  343. 11:14chat window up at the top, we have its
  344. 11:15response, but you'll notice that in
  345. 11:17between, we have this sort of like
  346. 11:18grayed-out section here. Okay, in this
  347. 11:20grayed-out section is the thinking that
  348. 11:22the model is doing before it gets back
  349. 11:24to us. And so, basically, you know, if
  350. 11:27this was ChatGPT back from 2022 or so,
  351. 11:30all we would have gotten is this. But
  352. 11:32because I'm telling it to take actions
  353. 11:34in the real world, it's capable of one,
  354. 11:37observing. And so, it observes all of
  355. 11:40this text and all of its reply as
  356. 11:42context.
  357. 11:44Two, thinking. So, it's capable of doing
  358. 11:46a bunch of thinking on what to do next.
  359. 11:49And then three, acting. And so then it's
  360. 11:51capable of saying, "Hmm, the user
  361. 11:52probably wants me to do some research. I
  362. 11:54have access to a few tools available.
  363. 11:56One of the tools lets me search the web.
  364. 11:58Let me pump in a search term." It then
  365. 12:00compiled all of this information, and
  366. 12:02then it just repeated the same thing. It
  367. 12:04then with all this context said, "Okay,
  368. 12:06I'm observing. Not only do I have these
  369. 12:07messages, but I also now have a bunch of
  370. 12:09research. Let me think about what to do
  371. 12:10next. Have I achieved the goal of the
  372. 12:13user compiling 10 plus empirical
  373. 12:14sources?" And you know, after it's made
  374. 12:16its sort of observation and thought on
  375. 12:18the reasoned about it, then it's
  376. 12:20deciding to act. And what it's ended up
  377. 12:21doing after 58 seconds is giving me this
  378. 12:23structured evidence report. So, this is
  379. 12:25an example of something that might have
  380. 12:26looped two times, three times, but the
  381. 12:28more intelligent and capable these
  382. 12:30models are getting, um the longer that
  383. 12:32they're running autonomously without us.
  384. 12:34Hopefully, this isn't rocket science to
  385. 12:35anybody here, but in a nutshell, this is
  386. 12:37more or less what's always occurring
  387. 12:39non-stop every time you talk to a model.
  388. 12:41With all that being said, let's really
  389. 12:43quickly cover how to set these different
  390. 12:44models up. I'm going to be using Codex,
  391. 12:47Claude Code, and Antigravity. You don't
  392. 12:50need to know anything about any of these
  393. 12:52platforms in order to run these
  394. 12:53examples. And if you're already very
  395. 12:55familiar with, let's say, I don't know,
  396. 12:56Claude Code, and you've chosen to use
  397. 12:58that as your main agentic coding
  398. 12:59platform moving forward, you can skip
  399. 13:01over to the next section of the video.
  400. 13:03But I want to make sure that we all have
  401. 13:04an equal playing ground here, and we all
  402. 13:06understand how each of these platforms
  403. 13:07work under the hood. So, there are three
  404. 13:09major platforms. The first is Codex,
  405. 13:11which is owned, managed, and run by
  406. 13:13OpenAI. The second is Claude Code, which
  407. 13:16is owned, managed, and run by Anthropic.
  408. 13:18And the third is Google's Antigravity,
  409. 13:20which as I'm sure you can imagine is
  410. 13:22owned, managed, and run by Google. In
  411. 13:24order to start with Codex, what you
  412. 13:26first have to do is sign up to an Open
  413. 13:28AI account. The way you do so is just
  414. 13:30look up Open AI on Google, get to a page
  415. 13:33that looks anything like this, and then
  416. 13:35just go to the top right-hand corner
  417. 13:36where it says try Chat GPT. After that,
  418. 13:38you'll be taken to a page that looks
  419. 13:39something like this. You can continue
  420. 13:41with Google, your phone, or whatever you
  421. 13:43want. And if you choose to chat with the
  422. 13:45model and then come back at any point in
  423. 13:47time, just head to the top right-hand
  424. 13:48corner for that model again. So, I'm
  425. 13:50going to pretend that I haven't made an
  426. 13:52account before and I'll continue with
  427. 13:53Google. After some brief onboarding
  428. 13:55instructions, you'll have access to a
  429. 13:56page like this. But, this is just Chat
  430. 13:59GPT, which is more akin to a chatbot
  431. 14:01than anything else. We want to take this
  432. 14:03to the AI agent world. And so, in order
  433. 14:05to do that, we need to use their
  434. 14:07dedicated AI agentic coding platform
  435. 14:09Codex. So, Googling Open AI's Codex or
  436. 14:12something like that will take you to a
  437. 14:13page that looks like this, and then you
  438. 14:15can just click download for macOS. By
  439. 14:17the way, I'm on a Mac, so that button's
  440. 14:19automatically going to pop up for me.
  441. 14:21But, the Codex app is now also available
  442. 14:23on Windows starting March 2024th and
  443. 14:25beyond. The way you install things on a
  444. 14:27Mac is you just take this window, drag
  445. 14:29Codex over to applications, and then
  446. 14:31you're done. Once you're inside, if you
  447. 14:32wanted to build a website or something,
  448. 14:34just head over to this middle, create a
  449. 14:36new folder, call it whatever you want.
  450. 14:38So, I'll just go to downloads and then
  451. 14:39go a new folder, example.
  452. 14:42Open it within it, and now you're inside
  453. 14:44of this folder. Here, you can ask the
  454. 14:45model to do whatever you want. And so,
  455. 14:47what I'm going to say is make a brief
  456. 14:48portfolio site about Nick Surive. Keep
  457. 14:51it super simple and minimal.
  458. 14:53It'll now do some thinking.
  459. 14:55In our case, I actually have a design
  460. 14:56taste front-end skill, which improves
  461. 14:58its ability to create like sleek,
  462. 15:00high-quality looking designs.
  463. 15:02And now, it's looking through my own
  464. 15:04workspace to put together this cool,
  465. 15:05sexy site for me. I'm also going to ask
  466. 15:08it to open it.
  467. 15:09Uh and the way that all AI agent
  468. 15:11platforms work now is you have the
  469. 15:13ability to put a
  470. 15:15queued message in, which you can also
  471. 15:17choose to send immediately via steer. In
  472. 15:20In case, I'll just wait until it's done.
  473. 15:21It'll consume this open it message and
  474. 15:23then it'll just open it for me in a new
  475. 15:25tab. Once it's done, the open it message
  476. 15:27will be fed in and it's just going to
  477. 15:28open this for me in a new tab. Now, I'm
  478. 15:31kind of zoomed in here, so if I zoom in
  479. 15:33a little bit more, you'll see that this
  480. 15:34is just a a simple one-page site that
  481. 15:36says Nick Sarif builds clear modern
  482. 15:38digital work. Here's some information
  483. 15:40about me and here's a contact page. Not
  484. 15:42rocket science, but this is how easy it
  485. 15:43is to like build web stuff. Claude is
  486. 15:45pretty similar. Just Google Claude sign
  487. 15:47up or something like that and you'll be
  488. 15:49taken to a page that looks like this.
  489. 15:51Here, you just enter your email address
  490. 15:52or in my case, continue with Google. In
  491. 15:54Claude's case, in order to use Claude
  492. 15:56code, you do have to pay for it. And so,
  493. 15:58there is a pro plan here that's $17 per
  494. 16:01month with an annual subscription or 20
  495. 16:03bucks if billed monthly.
  496. 16:05I'm not working for Claude or anything
  497. 16:07like that. I don't have any sort of
  498. 16:09affiliation with Anthropic in that way,
  499. 16:11but I will say that I received probably
  500. 16:14a 100 to 200 x return on my investment
  501. 16:17with an agent coding platform, whether
  502. 16:19it's Claude or whether it's Gemini or
  503. 16:21whether it's Codex. So, my
  504. 16:23recommendation for you, if this seems a
  505. 16:24little bit steep, is bite the bullet,
  506. 16:26pay it and learn whatever you can to
  507. 16:29make a return on investment with that
  508. 16:30money in the first month because this
  509. 16:32stuff is really quite powerful. Assuming
  510. 16:34you're done, just type Claude code
  511. 16:36desktop download or something like that.
  512. 16:38You'll be taken to a page that looks
  513. 16:39like this, which allow you to download
  514. 16:41it for Mac OS, Windows or even Windows
  515. 16:43ARM 64. So, I'm going to give my Mac OS
  516. 16:46thing a quick click. Then I'll go to the
  517. 16:48top right-hand corner. I'll just open
  518. 16:49Claude up just like I did with Codex.
  519. 16:51That'll take me to a page like this and
  520. 16:53then I just drag this over to the right.
  521. 16:54And then once you're done, you'll be
  522. 16:55taken to a chat page that looks
  523. 16:57something like this. What we really want
  524. 16:59is we want this code button, so I'm
  525. 17:00going to give that a click. Then here,
  526. 17:02all we need to do is just choose a
  527. 17:04folder to work in and then we can put in
  528. 17:06a quick request. So, I'm just going to
  529. 17:07choose a general folder Nick Sarif. Then
  530. 17:10I'm going to say bypass permissions,
  531. 17:12which might seem a little bit scary to
  532. 17:13you, but it just makes the model act
  533. 17:15independently. Then finally, I'm going
  534. 17:17to say, "Hey, make a brief portfolio
  535. 17:19site about Nick Sheraif. Super simple
  536. 17:21and minimal."
  537. 17:22And so, just like Codex designed it a
  538. 17:24moment ago with its various UX uh
  539. 17:27features, we have the same thing here
  540. 17:28with Claude Code. It's going to ask to
  541. 17:30access some files in my folder.
  542. 17:33And in addition to having the message
  543. 17:35box, we also have this sort of grayed
  544. 17:36out shining uh decal here, which is sort
  545. 17:39of it's like sinking, if you think about
  546. 17:41it, as well as its tool calls.
  547. 17:43And what it's going to do now is
  548. 17:44actually build me a brief little site.
  549. 17:46And then just like I did before, I'll
  550. 17:48just say, "Open it."
  551. 17:50That's going to queue it, and now I can
  552. 17:52have a conversation with Claude. And now
  553. 17:54we have the actual portfolio, which as
  554. 17:55you guys can see here is done in
  555. 17:57significantly more minimal fashion,
  556. 17:59okay? So, this is Nick Sheraif, builder
  557. 18:01automation expert software engineer.
  558. 18:02Now, unlike with ChatGPT and then Claude
  559. 18:06for Anti-Gravity, odds are you probably
  560. 18:08already have like a Google or a Gmail
  561. 18:10account set up. So, all you have to do
  562. 18:11is just look up Google Anti-Gravity
  563. 18:13download, then click download for Mac
  564. 18:15OS. In my case, I have Apple silicon on
  565. 18:18Mac. If you guys don't know what you
  566. 18:20have, just type about this Mac, and then
  567. 18:22if it says Intel up here and chip,
  568. 18:23you're an Intel. If it's a M something,
  569. 18:26then you're Apple silicon. And you can
  570. 18:27do something similar for Windows and
  571. 18:29Linux, as well. And once I give that a
  572. 18:30click, we'll be taken to a very
  573. 18:32similar-looking page here, and then I
  574. 18:33can just drag Anti-Gravity over to
  575. 18:35applications. The very first time you
  576. 18:36open up Anti-Gravity, it'll look
  577. 18:38something like this. In your case, maybe
  578. 18:39it'll be dark mode, or maybe it'll be
  579. 18:41entirely light. I just have some styling
  580. 18:43settings, which is why mine might look a
  581. 18:45little different from yours. You may
  582. 18:46also have to log in, unless Google
  583. 18:48logged you in automatically. In my case,
  584. 18:50it logged me in automatically because
  585. 18:51I've used it before. Assuming that
  586. 18:53you've done that though, on the
  587. 18:54right-hand side, you'll see an agent
  588. 18:56model. And this agent model is very
  589. 18:58similar to what we saw with Codex and
  590. 18:59then Claude Code. All we have to do is
  591. 19:01just ask it to make a brief portfolio
  592. 19:03site about Nick Sheraif. You'll see here
  593. 19:04that the UX is just a little bit
  594. 19:06different, right? We have a little
  595. 19:07generating tab down here. Obviously, we
  596. 19:09have uh multiple settings with fast and
  597. 19:12Gemini 3.1 Pro. We have this little
  598. 19:14thinking tab. Uh it tells you how long
  599. 19:16it's been doing it. If it has to do any
  600. 19:18web searches, it does so over here.
  601. 19:20Hopefully you guys are seeing these are
  602. 19:22all just flavors that are slightly
  603. 19:24different, but ultimately are the same
  604. 19:26thing. I'm just going to write open it.
  605. 19:27That'll be added as a pending message,
  606. 19:29and then it'll open this up in a browser
  607. 19:31tab. As you see here, Gemini produced
  608. 19:33what I would probably consider to be the
  609. 19:34sexiest of all websites, which makes
  610. 19:36sense. Uh one thing I'll talk about in a
  611. 19:38moment is how much better it is at
  612. 19:40front-end design and so on and so forth.
  613. 19:42And yeah, we have a very simple and and
  614. 19:43straightforward site here. So, um this
  615. 19:45links to all of my resources, left
  616. 19:47click, YouTube, and so on and so forth.
  617. 19:49I probably like this one the best. From
  618. 19:51here on out, most of the conversations
  619. 19:53and the user experiences are going to be
  620. 19:55really similar between the agent coding
  621. 19:57platforms. So, while I am going to use
  622. 19:59multiple just to show you guys how some
  623. 20:01of their quirks interact, uh for the
  624. 20:03most part, I want you guys to know that
  625. 20:04the UXs are are very very similar these
  626. 20:07days. Like the thinking tabs, they're
  627. 20:08going to be the same. Some people will
  628. 20:10probably say that there are slight
  629. 20:11differences between them and so on and
  630. 20:13so forth. For instance, I'm a big fan of
  631. 20:15the little Space Invader icon in that
  632. 20:16Claude Code has. Uh but for all intents
  633. 20:18and purposes, I'm just going to assume
  634. 20:20that you're picking up the UX here as
  635. 20:22you use these models, and focus less on
  636. 20:24like the tiny little stuff and more on
  637. 20:26how to orchestrate and then prompt these
  638. 20:28for higher quality responses. If you
  639. 20:29guys want to see like step-by-step
  640. 20:31walk-throughs of these platforms, I'm
  641. 20:33going to put some little links up above
  642. 20:35my left shoulder here, and you can uh
  643. 20:37click on them anytime to go learn that
  644. 20:38sort of stuff. Next up, I want to talk
  645. 20:40about what makes these AI coding
  646. 20:41platforms different from one another.
  647. 20:44Not on a user experience um angle, but
  648. 20:47from an intelligence angle, from a what
  649. 20:49they could do angle as well. So, as you
  650. 20:51saw there, there were three different
  651. 20:53models. There was Claude, which was
  652. 20:55wrapped around Claude Code, Gemini,
  653. 20:58which was wrapped around antigravity,
  654. 21:00and then GPT, in my case 5.4, which is
  655. 21:03wrapped around Codex.
  656. 21:05And I think that each of these models
  657. 21:07are really similar at this point in
  658. 21:08intelligence-wise, but there are some
  659. 21:10pros and cons to each that basically
  660. 21:13like improve how they perform by a few
  661. 21:15percentage points. So, Claude might be,
  662. 21:17you know, 2% better at these, you know,
  663. 21:19Gemini might be 5% better at these, GPT
  664. 21:22might be 1% better than these. I'm just
  665. 21:24pulling out numbers out of my butt. But,
  666. 21:26I'm making them really small because I
  667. 21:27do want to really drive home the point
  668. 21:29that these models are so gosh darn
  669. 21:31intelligent these days that these minor
  670. 21:33differences only make sense at the
  671. 21:34bleeding edge and at the frontier. For
  672. 21:36most purposes, either of these are going
  673. 21:39to be sufficient. So, Claude has the
  674. 21:41most interpretable reasoning. You
  675. 21:43remember how I could click open that
  676. 21:45little reasoning tab a moment ago? Well,
  677. 21:47at least as of the time of this
  678. 21:48recording, Claude is incredible at
  679. 21:50making that reasoning tab really, really
  680. 21:52interpretable. You know exactly what
  681. 21:54Claude is doing at basically every step
  682. 21:55of the process when you use Claude code
  683. 21:58to visualize that reasoning. And that
  684. 21:59makes it really good for orchestration
  685. 22:02and then agentic workflows because you
  686. 22:05can see the decisions of the model is
  687. 22:06making in real time. And in doing so,
  688. 22:08you can also steer the model, stop the
  689. 22:10model, pause it, or give it new
  690. 22:12resources halfway through. I can't say
  691. 22:14the same about both Gemini and GPT. I
  692. 22:16think they're a lot less interpretable
  693. 22:17and it's a lot less accountable. You
  694. 22:19know, Claude is sort of a partner that
  695. 22:21you build things with along the way,
  696. 22:23whereas Gemini and GPT are almost just
  697. 22:24like, I don't know, they're missiles.
  698. 22:26You set your target, you click the
  699. 22:27button, and then they go. Now, there are
  700. 22:29some cons. Claude is a little bit slower
  701. 22:31unless you use fast mode, which is what
  702. 22:34I tend to use, although keep in mind
  703. 22:35that'll burn a ton of credits. And then
  704. 22:37I find that it's weaker at front end or
  705. 22:39design than a model like Gemini.
  706. 22:41Gemini is really good at design and
  707. 22:43front ends. As you guys just saw a
  708. 22:45moment ago, Claude picked a really
  709. 22:46minimalistic sleek theme. Gemini did
  710. 22:49some upscale stuff that still looked
  711. 22:51sleek, clean, but had like that
  712. 22:53isomorphic glass. And then GPT, maybe
  713. 22:56because of my design taste scale or
  714. 22:57something else, was kind of like more
  715. 22:59complex and had uh a little bit clunkier
  716. 23:01of a design. Well, in general, I find
  717. 23:03that this pattern remains the same.
  718. 23:05Anytime I want to design a really clean
  719. 23:07front end, I'm going to use Gemini for
  720. 23:08that. It's also got superior multimodal
  721. 23:10abilities. That just means there's
  722. 23:12actual like endpoints using the Gemini
  723. 23:14API um where it can understand video.
  724. 23:17Right now, Claude and GPT both really
  725. 23:18struggle with this, although you can
  726. 23:20build custom pipelines to do that, which
  727. 23:21I should have showed you guys about. It
  728. 23:22also has the ability to use a fast
  729. 23:24output, which means it writes really,
  730. 23:26really quickly if need be, um but they
  731. 23:27don't have access to a dedicated fast
  732. 23:30mode where you could pay more money to
  733. 23:31use them really quick. I think it's the
  734. 23:33least interpretable of the models, and
  735. 23:34personally I find the quality is quite
  736. 23:36inconsistent. There's some days when
  737. 23:38I'll prompt it and it'll do quite
  738. 23:39incredible, then other days where I'll
  739. 23:40prompt it and it will just absolutely
  740. 23:42crap the bed.
  741. 23:43You know, at least Claude's quite
  742. 23:44consistent in that way, despite the fact
  743. 23:46that maybe it's a little bit worse at a
  744. 23:47few things.
  745. 23:48Finally, there's GPT. There's the Codex
  746. 23:51series of models, the 5.4 series of
  747. 23:53models now. These are the best at
  748. 23:54back-end programming. I think they're
  749. 23:56also the best at like um absolute
  750. 23:58mathematics, which probably feeds into
  751. 24:00that. They're really great at
  752. 24:01test-driven development, and you know
  753. 24:03how I mentioned earlier Gemini and GPT
  754. 24:05are more like rockets that you point at
  755. 24:06a at a at a place and then they go. Um
  756. 24:08well, these test-driven development
  757. 24:11approaches essentially mean you just
  758. 24:12outline that definition of done, and
  759. 24:14then it fires and just goes autonomously
  760. 24:16until it reaches that. There's also
  761. 24:18quite a big ecosystem of different apps,
  762. 24:19and you know, there's a lot of um
  763. 24:21documentation online about how to use
  764. 24:23various GPT workflows and stuff like
  765. 24:25that, because this was the first major
  766. 24:27player to the AI agent market. I'd give
  767. 24:30it sort of like a uh you know, two out
  768. 24:32of three on the rest of these. I think
  769. 24:33Claude is much better at its
  770. 24:35interpretability, it's much better at
  771. 24:36orchestration and stuff like that. But
  772. 24:38GPT, being a model that just came out
  773. 24:40quite recently, a 5.4 anyway, is
  774. 24:43obviously sort of like topping the
  775. 24:44charts right now on a lot of stuff. Just
  776. 24:45some caveats there, a lot of people
  777. 24:47treat this as like
  778. 24:49>> [snorts]
  779. 24:49>> anathema for you to claim that, you
  780. 24:51know, Claude is better than GPT at this
  781. 24:53thing, and Gemini is better than than
  782. 24:55Claude at that thing. The reality is, as
  783. 24:57I mentioned and alluded to at the
  784. 24:59beginning, there are very minor
  785. 25:00differences between these models at this
  786. 25:02point. All of them are basically trained
  787. 25:03on the entirety of the internet as is.
  788. 25:05And so because of this, the slight
  789. 25:08differences in capabilities in the model
  790. 25:10tend to have more to do with like when
  791. 25:11they were trained and how recent it is
  792. 25:14versus, you know, some inherent like
  793. 25:15cool new design technique. Really,
  794. 25:17they're just training these galaxy-sized
  795. 25:19brains on the entire internet at this
  796. 25:21point. So because we're talking about
  797. 25:22the LLM intelligences, you know, if like
  798. 25:24GPT was trained after Claude, GPT's
  799. 25:26probably going to be a little bit better
  800. 25:27in certain circumstances. If Gemini's
  801. 25:29trained after GPT, it'll be better. But
  802. 25:32all that stuff resets with the next
  803. 25:33generation. So though I am going to be
  804. 25:35showing you guys some cool multi-MCP
  805. 25:37orchestration uh techniques later on, I
  806. 25:39want you to know that you don't have to
  807. 25:40treat all this super seriously. You can
  808. 25:41also just pick one model and then use
  809. 25:43that. Okay, next up I want to chat
  810. 25:45agents.md and then how to build a
  811. 25:47self-modifying and self-correcting
  812. 25:49system prompt that significantly
  813. 25:51minimizes the number of errors that you
  814. 25:53get as you build things with these AI
  815. 25:55agents. So for the purposes of this
  816. 25:57demonstration, I'm going to be using
  817. 25:58antigravity and through it the Gemini
  818. 26:00series of models. When you open up
  819. 26:02antigravity, you have a little window
  820. 26:03that looks like this. Generally, I
  821. 26:05divide this into three panes. You have
  822. 26:07your explorer on the left-hand side,
  823. 26:08your file editor in the middle, and then
  824. 26:10you have your agent on the right. What
  825. 26:12I'm going to do for the purposes of this
  826. 26:13demo is I'll just click open folder and
  827. 26:15then I'm going to go to antigravity
  828. 26:17example and just open this up.
  829. 26:19Okay, and what I want to do here is I
  830. 26:20just want to show you how all of this
  831. 26:21stuff works to start.
  832. 26:23As you guys could see on the left-hand
  833. 26:24side, we have a file called Gemini.md.
  834. 26:27Now what occurs is when you talk to this
  835. 26:29model over here, hey, what's up?
  836. 26:32Basically, what's occurring is this file
  837. 26:35is being prepended to the very top of a
  838. 26:38conversation chain. And so if I open up
  839. 26:40this file right now, you see how it's
  840. 26:41empty, there's nothing in it. Well, when
  841. 26:43I started this conversation and said,
  842. 26:45"Hey, what's up?" Okay, it knows that my
  843. 26:47name is Nick, but it does it knows this
  844. 26:49because of the fact that I'm signed in
  845. 26:50as Nick Surave.
  846. 26:52Now I want you to see what happens if I
  847. 26:54paste in my name is Antonio Banderas,
  848. 26:56refer to me as such, always always also
  849. 26:58always sign off super kawaii desu. So,
  850. 27:00I'm going to go here to the top right
  851. 27:01hand corner and I'll say, "Hey, what's
  852. 27:04up?"
  853. 27:05And after initializing a new model,
  854. 27:09notice how it's now going to return
  855. 27:11something quite different to what we had
  856. 27:13a moment ago. The reason why is of
  857. 27:15course this gemini.md is just a
  858. 27:17templated structured prompt that is
  859. 27:20basically always inserted into the
  860. 27:22beginning. Okay? The same thing applies
  861. 27:24with Codex, the same thing applies with
  862. 27:26Claude Code.
  863. 27:27But the names of the files are a little
  864. 27:29bit different. So, if I was in, let's
  865. 27:31say, Codex for instance, I wouldn't call
  866. 27:32this a gemini.md, I'd call this an
  867. 27:34agents.md. If I was in Claude Code, I
  868. 27:36wouldn't call this an agents.md, I'd
  869. 27:38call this a Claude.md. Whatever file you
  870. 27:41use here doesn't really change the idea.
  871. 27:43The idea is that at the very top of any
  872. 27:45prompt, you just have this file
  873. 27:47prepended to it.
  874. 27:48The reason why this is so powerful is
  875. 27:50because you now have the ability to
  876. 27:52statically template out the same prompt
  877. 27:55over and over and over again on every
  878. 27:57independent session. This may seem like,
  879. 27:59well, why don't you just copy and paste
  880. 28:01the same thing in instead of having to
  881. 28:02use this elaborate file system
  882. 28:03structure? And the reason why is because
  883. 28:05what you can do is at the very beginning
  884. 28:07of this file, you can actually contain
  885. 28:10within it like a list of lessons or
  886. 28:12learnings from previous instances.
  887. 28:15Then you can build in a like a meta
  888. 28:16prompt structure where before a model
  889. 28:18signs off, before it finishes whatever
  890. 28:20it's doing, it always updates that file
  891. 28:22with more and more and more knowledge.
  892. 28:24In that way, okay, you can build a
  893. 28:26high-quality list of like memories,
  894. 28:28preferences, and rules, not to mention
  895. 28:31things to avoid, that significantly
  896. 28:33improves your agent's ability to operate
  897. 28:35over a long time scale. And just to show
  898. 28:37you guys what I mean, let me show you a
  899. 28:38diagram. In this hypothetical instance,
  900. 28:41we're going to be using gemini.md.
  901. 28:43And basically what will occur every time
  902. 28:44is a new session is going to start over
  903. 28:46here. The agent will first read
  904. 28:48Gemini.md.
  905. 28:50You'll then give it a task like, "Hey,
  906. 28:52build me a website that does whatever."
  907. 28:55Now, it'll return the website for me,
  908. 28:57and then I'll say, "I don't like this.
  909. 28:59No dark mode."
  910. 29:01After I give it its feedback of no dark
  911. 29:03mode, rather than just correcting the
  912. 29:05build, it'll actually write that to my
  913. 29:07Gemini.md for next time, which allow the
  914. 29:10agent to continue working with the rule
  915. 29:11applied. When the session ends and a new
  916. 29:13session starts, now the agent will read
  917. 29:16the Gemini MD, but the Gemini.md will
  918. 29:18have an additional rule placed, okay?
  919. 29:20This is my file over here. It'll say,
  920. 29:22"No dark mode."
  921. 29:23And that means the next time I ask it to
  922. 29:25build me a website or any sort of web
  923. 29:26property, it'll see no dark mode, and
  924. 29:28then it won't make that mistake again.
  925. 29:29This lets your knowledge accumulate over
  926. 29:32sessions. The first time that you use,
  927. 29:34you know, Gemini or Claude Code or or
  928. 29:36Codex or whatever, you know, you're only
  929. 29:38going to have, let's say, one rule or
  930. 29:40one preference stored. And so, the
  931. 29:42number of errors that the model makes,
  932. 29:44errors relative like your preferences,
  933. 29:45will be pretty high. The second time
  934. 29:48that you use it, though,
  935. 29:49the number of errors or issues that it
  936. 29:51makes that don't line up with your
  937. 29:52preferences will go down.
  938. 29:54The third time, they'll go down further.
  939. 29:57The fourth time, it'll go down further.
  940. 29:59And the fifth time, it'll go really,
  941. 30:01really low, to the point where it maybe
  942. 30:02it makes zero errors at all.
  943. 30:04You can see that um sort of
  944. 30:05diagrammatically over here, with when
  945. 30:07you start, your thing has zero rules,
  946. 30:08okay? As it grows longer and longer and
  947. 30:11longer, you're writing more and more and
  948. 30:13more and more rules. Um the agents get
  949. 30:15better and better and better at
  950. 30:16understanding and then um
  951. 30:18anticipating as well your preferences.
  952. 30:20So, what does this actually look like in
  953. 30:21practice? Well, it's not all that
  954. 30:23difficult, and you can just append or
  955. 30:25prepend this to any Gemini, Claude, or
  956. 30:28agent's MD, however you like. It also
  957. 30:30doesn't need to be this long, although I
  958. 30:31did want to go into a fair amount of
  959. 30:32detail here with you. So, you can
  960. 30:34absolutely just turn this into like a I
  961. 30:35don't know, a three or four-line
  962. 30:37snippet.
  963. 30:38Essentially, before we start any task,
  964. 30:40read this entire file.
  965. 30:41This file contains a growing rule set
  966. 30:43that improves over time. At session
  967. 30:45start, I want you to read the entire
  968. 30:47learned rule section before doing
  969. 30:48anything.
  970. 30:49How it works. When the user corrects you
  971. 30:51or you make a mistake, immediately
  972. 30:52append a new rule to the learned rule
  973. 30:54section at the bottom of this file.
  974. 30:57Rules are numbered sequentially and
  975. 30:58written as clear imperative
  976. 30:59instructions. The format is category
  977. 31:02never or always do X because Y, and then
  978. 31:05here's some more formatting
  979. 31:06instructions.
  980. 31:07When do you add a rule? Add a rule when
  981. 31:09the user explicitly corrects your
  982. 31:10output. When the user rejects a file
  983. 31:12approach or pattern. When you hit a bug
  984. 31:14caused by a wrong assumption or when the
  985. 31:16user states a preference. Okay, and then
  986. 31:18it'll give some examples here of
  987. 31:19different rules and code. Then we have
  988. 31:20the learned rules down here. So, what
  989. 31:22I'll do, just to show you guys what this
  990. 31:24looks like, is I'll say, "Build me
  991. 31:27a simple portfolio site for Nick Saraf."
  992. 31:30And I'm going to have it go accomplish a
  993. 31:31task for me. And then, I'm inherently
  994. 31:33and intentionally going to give it some
  995. 31:36instructions.
  996. 31:37You see, the very first thing it did was
  997. 31:38analyze the gemini.md. And so, now it
  998. 31:41actually has this entire file as context
  999. 31:43inside of its thread. You can't see that
  1000. 31:46context here because obviously they
  1001. 31:47don't want to just muck up your your
  1002. 31:49conversation thread, but it is literally
  1003. 31:51like if you just pasted this entire
  1004. 31:52thing directly in, okay? So, it's going
  1005. 31:55to be reading that constantly as it's
  1006. 31:56building up the rest of our website.
  1007. 31:58And you can see that it's like it's
  1008. 31:59built some cool terminal display here.
  1009. 32:02It's using a library called Vit, which
  1010. 32:03is probably like the best front-end
  1011. 32:05library. Let's see what it does. Okay,
  1012. 32:07this website is looking really, really
  1013. 32:09sexy, super clean, and it clearly went
  1014. 32:10above and beyond with my spec. However,
  1015. 32:13I don't like how it's dark mode. So,
  1016. 32:14what I'm going to do is go back here and
  1017. 32:16then give it some instructions. "Quit
  1018. 32:18doing things in dark mode."
  1019. 32:21And the idea here is, when I give it an
  1020. 32:23instruction like quit doing things in
  1021. 32:25dark mode, what it's going to do is it's
  1022. 32:26going to take my message and then say,
  1023. 32:29"Hey, let's update our gemini.md to
  1024. 32:32never create applications in dark mode.
  1025. 32:34It's a user preference."
  1026. 32:36If I scroll down here now, you can
  1027. 32:38actually see that this style has been
  1028. 32:40added. And so, if the next time I run a
  1029. 32:43model
  1030. 32:44and instantiate anti-gravity, I say,
  1031. 32:46"Hey, I'd like you to build me a
  1032. 32:46website." You'll actually have this up
  1033. 32:48at the very, very top of its prompt.
  1034. 32:51Meaning that I'm never, ever going to
  1035. 32:52have a dark mode website again.
  1036. 32:54In this way, this will continuously get
  1037. 32:56closer and closer to my preferences
  1038. 32:58until the number of rules becomes so
  1039. 33:00exhaustive that, you know, it'd actually
  1040. 33:01be counterproductive. In practice, I
  1041. 33:04haven't actually hit this limit yet. I
  1042. 33:05think this just gets better and better
  1043. 33:06and better over time, but I could
  1044. 33:08hypothetically see if you were to get to
  1045. 33:09a point where there's a thousand
  1046. 33:10independent rules, some of them would
  1047. 33:11probably start stepping on its its toes.
  1048. 33:14Um this sort of self-modifying Claude
  1049. 33:16agents or Gemini.md is a very, very high
  1050. 33:19ROI design pattern. So, whatever you're
  1051. 33:21building with an AI agent, whether
  1052. 33:22you're using them for business,
  1053. 33:23personal, or programming tasks, I would
  1054. 33:25always recommend to have something like
  1055. 33:26this in your directory. And as you can
  1056. 33:28see, it's now modified the site. We
  1057. 33:29don't actually have that anymore. A lot
  1058. 33:31cleaner, and it also fixed up the images
  1059. 33:33and made it look really sexy. The way
  1060. 33:34this works is at the very top level, we
  1061. 33:36have a global Claude agents or
  1062. 33:39Gemini.md. And these are user-wide rules
  1063. 33:42that apply to all of the projects that
  1064. 33:44you start. And so, the very top, you'll
  1065. 33:46have this sort of injected, and you can
  1066. 33:48set this using a variety of different
  1067. 33:50formatting conventions and stuff. You
  1068. 33:52could look it up for the specific uh
  1069. 33:54agent platform that you're using. And if
  1070. 33:56you're doing Claude or something like
  1071. 33:57that, it's going to be stored in a a
  1072. 33:59tilde. dot Claude {slash} and then there
  1073. 34:03are a variety of other conventions
  1074. 34:04regardless of whatever platform you're
  1075. 34:06using that you guys can also After it's
  1076. 34:08injected the global agents.md, it'll
  1077. 34:11then inject the local Claude.md. And so,
  1078. 34:14what you could do is you could have a
  1079. 34:15global Claude.md, okay, that has
  1080. 34:19wide-ranging user preferences updated,
  1081. 34:21and then a local project.md that has
  1082. 34:24specific project preferences updated.
  1083. 34:26And then underneath, you also have uh
  1084. 34:27skills, and then you're finally in-line
  1085. 34:30prompt. And I'll touch on the skill
  1086. 34:32section in a moment. But in that way you
  1087. 34:34can collapse a ton of context and a ton
  1088. 34:36of sort of functionality into very few
  1089. 34:39tokens, which is important because your
  1090. 34:40bill both per token and then the quality
  1091. 34:42of the models tend to degrade the longer
  1092. 34:44the token context windows get. Next up I
  1093. 34:46want to talk a little bit about agent
  1094. 34:48skills. And this isn't going to be an
  1095. 34:49exhaustive resource. If you guys want a
  1096. 34:51super in-depth way to look at skills,
  1097. 34:53definitely just check out my full
  1098. 34:54end-to-end Claude code skills course.
  1099. 34:57But agent skills, for those of you guys
  1100. 34:59that don't know, is just a simple
  1101. 35:01repeatable way that you can standardize
  1102. 35:04workflows.
  1103. 35:05Now, this is important because large
  1104. 35:07language models are very flexible. So,
  1105. 35:09if you give them a non-super tightly
  1106. 35:11scoped task, they'll tend to produce a
  1107. 35:13variety of different results for you.
  1108. 35:15Well, skills are just a way of basically
  1109. 35:17turning that whole, you know, vagueness,
  1110. 35:20that whole statistical variance into
  1111. 35:22like a really straight-line
  1112. 35:24deterministic path where it just does
  1113. 35:25the same thing over and over and over
  1114. 35:27and over and over again.
  1115. 35:28And so, skills are offered now on all
  1116. 35:30major platforms. We've all adopted them.
  1117. 35:32So, you have Codex skills, you have
  1118. 35:34Gemini skills, and then you also have
  1119. 35:37Claude code skills. And they have very
  1120. 35:39particular specs and they look really,
  1121. 35:41really similar to one another. So, it's
  1122. 35:42worth me at least going over to
  1123. 35:43high-level what they look like. To make
  1124. 35:45a long story short, these are just files
  1125. 35:47that will exist somewhere within our
  1126. 35:49workspace. These files will have sort of
  1127. 35:51this little title section up here, which
  1128. 35:53you know is a title because there'll be
  1129. 35:54three hyphens at the top and three
  1130. 35:56hyphens at the bottom. Inside of the
  1131. 35:58file you can give it a name like PDF
  1132. 35:59processing, a description like extract
  1133. 36:01text and tables from PDFs,
  1134. 36:04and then you can even do licenses and
  1135. 36:05metadata and so on and so forth. I don't
  1136. 36:07actually do any of this stuff. My skills
  1137. 36:09are almost always just name,
  1138. 36:10description, and then maybe some
  1139. 36:11optional
  1140. 36:13tools that it could use as well. Okay,
  1141. 36:15so I just want to give you guys a couple
  1142. 36:16of brief examples. I'm just going to go
  1143. 36:17over to Anthropic skills because they
  1144. 36:20have a a bunch of simple ones here that
  1145. 36:22we can use just to gain some context.
  1146. 36:24I'm going to go over to the skills
  1147. 36:25folder here and then click on I don't
  1148. 36:27know let's do algorithmic art.
  1149. 36:29We'll go skill.md cuz that's the file
  1150. 36:32and as you guys could see here we have
  1151. 36:34if I click on the raw you guys will see
  1152. 36:36we have the exact same format that I
  1153. 36:37showed you guys earlier. So this is a
  1154. 36:39skill that creates algorithmic art using
  1155. 36:41a particular library and what's cool is
  1156. 36:43it basically guides the model through
  1157. 36:45the same thing every time to get very
  1158. 36:47very similar algorithmic art generated.
  1159. 36:50You can see this is a pretty long skill
  1160. 36:51there's a lot going on right? So what
  1161. 36:53I'm going to do is I'm just going to
  1162. 36:54copy this whole thing and show you guys
  1163. 36:55how this works. In this way we can copy
  1164. 36:57and paste different standard operating
  1165. 36:59procedures to different models and then
  1166. 37:01get high quality results. So I'm going
  1167. 37:03to go over here and then you know just
  1168. 37:05because this is a one-shot prompt I'm
  1169. 37:07just going to feed all this in
  1170. 37:09and then I'm going to have this model
  1171. 37:10actually create things according to the
  1172. 37:12skill spec. So it's doing some thinking
  1173. 37:14and now it's asking me what do we want
  1174. 37:15to do with it and I'm going to say yes
  1175. 37:17save as skill
  1176. 37:19then run. And then I'm going to actually
  1177. 37:21have this like produce some sort of cool
  1178. 37:23algorithmic art. Now there's no template
  1179. 37:25file or anything like that so it's
  1180. 37:27actually going to go through the whole
  1181. 37:27process. It's going to create both the
  1182. 37:29skill directory which we can find right
  1183. 37:30over here now called algorithmic art
  1184. 37:33and then it's also going to create like
  1185. 37:34templates and a bunch of other stuff as
  1186. 37:36well. Okay and our algorithmic art flow
  1187. 37:38is just finished up so I'm actually just
  1188. 37:40going to open this so I can take a look
  1189. 37:41at it myself.
  1190. 37:43And we have it. There it is. This is now
  1191. 37:45creating algorithmic art as you guys
  1192. 37:47could see we have particles and so on
  1193. 37:48and so forth. I'm just going to
  1194. 37:50significantly decrease the number of
  1195. 37:51particles
  1196. 37:52maybe change the noise scale and the
  1197. 37:54turbulence. Actually move this around
  1198. 37:56and as you guys can see we we we are
  1199. 37:57actually producing a tremendous number
  1200. 37:59of particles here. This is this is
  1201. 38:00actually like rendering them directly in
  1202. 38:01my browser which is nuts.
  1203. 38:03Um so this is indeed algorithmic art.
  1204. 38:05It's it's really cool super sexy. I'm a
  1205. 38:07big fan. I don't know I mean it looks
  1206. 38:08kind of like hair but what are you going
  1207. 38:10to do?
  1208. 38:10I'm just going to regenerate a bunch
  1209. 38:12maybe change the accent colors. Okay
  1210. 38:14maybe we'll have this as my accent now
  1211. 38:16blue and then the background will be
  1212. 38:17kind of this and
  1213. 38:19I don't know my cool accent will be kind
  1214. 38:20of like this.
  1215. 38:22There you go. That looks pretty nice.
  1216. 38:24We can now kind of just create new ones
  1217. 38:27as we want and then we can also just
  1218. 38:28completely randomize them over and over
  1219. 38:30and over and over and over again. And
  1220. 38:31you can see it's actually still doing
  1221. 38:33some design in the background as we go.
  1222. 38:34So I'm just going to change the number
  1223. 38:36of particles to really low and then I'll
  1224. 38:38just redesign this over and over and
  1225. 38:39over and over again.
  1226. 38:42And I should note that like this is not
  1227. 38:43like a you know, it's not a piece of
  1228. 38:45software I downloaded. We actually just
  1229. 38:46built this. It's just we built this in a
  1230. 38:48much more standardized and you know,
  1231. 38:50consistent way which is really cool. So
  1232. 38:52obviously that's that's what I want. I
  1233. 38:53want the ability to share like
  1234. 38:55repeatable workflows where my agent can
  1235. 38:58build things that other people have
  1236. 38:59validated without me necessarily having
  1237. 39:01just to like copy and paste a piece of
  1238. 39:02software into my computer. Now remember
  1239. 39:04earlier how I said some models are
  1240. 39:06better at things than others and these
  1241. 39:08few percentage point differences can
  1242. 39:10make a lot of impact at the bleeding
  1243. 39:12edge or the frontier. Assuming you guys
  1244. 39:14are at the bleeding edge and the
  1245. 39:15frontier and those percentage point
  1246. 39:17differences stack up, then multi-agent
  1247. 39:20MCP orchestration is the pattern for
  1248. 39:23you.
  1249. 39:23Basically here what happens is you let
  1250. 39:26one model type be the manager or the
  1251. 39:29orchestrator. And that orchestrator will
  1252. 39:31take a task and then dole it out, okay,
  1253. 39:34and delegate sub chunks of that task to
  1254. 39:37different models. And so what's
  1255. 39:38occurring here is in this hypothetical
  1256. 39:40example we're using Claude code to be
  1257. 39:42our manager. We then give it some task
  1258. 39:44like, "Hey,
  1259. 39:46make me [snorts] a SaaS app that does X,
  1260. 39:50Y, and Z."
  1261. 39:51And then what it's doing is it's taking
  1262. 39:52my command and then splitting it into a
  1263. 39:54variety of different functions. There's
  1264. 39:56a front end task which is delegating to
  1265. 39:58Gemini to build the UI. There's a back
  1266. 40:01end task which is delegating to Codex to
  1267. 40:04build the API. There'll be some testing
  1268. 40:06that we need to occur
  1269. 40:07that we need to do which it'll delegate
  1270. 40:09to Codex to do the testing. Then finally
  1271. 40:11at the end we have Claude which will
  1272. 40:13collect and then validate the results.
  1273. 40:15And then if there are any discrepancies
  1274. 40:17or issues there, you know, we can loop
  1275. 40:18that back around hypothetically to
  1276. 40:20different models as we will.
  1277. 40:23And so this is a little bit more of an
  1278. 40:25advanced design pattern, and I don't
  1279. 40:26necessarily recommend you guys sign up
  1280. 40:28to a bajillion patterns and waste your
  1281. 40:30tokens that way unless you have to, but
  1282. 40:32I wanted to cover it because this is
  1283. 40:33sort of like the next generation of
  1284. 40:35model intelligence. It's where instead
  1285. 40:37of just sticking with one, you're
  1286. 40:38constantly querying different models for
  1287. 40:40things that they're a little bit better
  1288. 40:41at. All of this depends on this idea of
  1289. 40:44a router.
  1290. 40:45And so this router is more or less like
  1291. 40:47a decision hub or like a nexus.
  1292. 40:50When you give it a task or you give it
  1293. 40:52some sort of input, what it'll do is
  1294. 40:54it'll just divide it into different
  1295. 40:56subtasks that different models are
  1296. 40:58better than other models at. So for
  1297. 41:00instance, if we have like a high-level
  1298. 41:02task that has to do with replicating a
  1299. 41:04specific SaaS app, you know, and the the
  1300. 41:07model has decided that there's some
  1301. 41:09footage on the internet out there that
  1302. 41:11talks about how to build it, it'll
  1303. 41:12actually go delegate the video watching
  1304. 41:14step over to Gemini cuz Gemini's better
  1305. 41:16at multimodality and their endpoints
  1306. 41:18have built-in video understanding.
  1307. 41:20You know, if it identifies that we need
  1308. 41:22something with a lot of complex
  1309. 41:23reasoning, it'll route that over to
  1310. 41:25Claude. And if it identifies that we
  1311. 41:26need some form of sandboxed cloud code
  1312. 41:29execution, it'll do that in Codex cuz
  1313. 41:31they include that built-in. And maybe,
  1314. 41:32you know, I just wanted to show you guys
  1315. 41:34what an example would look like if you
  1316. 41:35had something that was outside of the
  1317. 41:36three. If you need real-time web data,
  1318. 41:38it might do that with Perplexity or
  1319. 41:40Perplexity's computer or something.
  1320. 41:42And what happens is, you know, we build
  1321. 41:43it all by parallelizing this big sweep,
  1322. 41:47and then at the very end we combine it
  1323. 41:48again with this router, which is
  1324. 41:50probably, you know, at least in my case
  1325. 41:52almost always going to be Claude Opus
  1326. 41:544.6, 4.7 by the time you guys are
  1327. 41:56reading it, and then that's what
  1328. 41:58ultimately unifies it before maybe doing
  1329. 42:00some additional Q&A, bug fixes, and
  1330. 42:02agent review, which I'll talk about
  1331. 42:03later. Now all of this sounds pretty
  1332. 42:05abstract, and you're like, "Okay, why
  1333. 42:06don't I just have all of this done in
  1334. 42:08one thread?" So let me show you a
  1335. 42:09practical way to actually do it. By the
  1336. 42:11way, all the files for this course you
  1337. 42:12can find in the top link in the
  1338. 42:14description below. What I'm going to do
  1339. 42:15is go back to Claude Code and open up a
  1340. 42:18new session. And then I'm going to
  1341. 42:20select this folder that I've actually
  1342. 42:21already created for this purpose called
  1343. 42:23multi-platform orchestration.
  1344. 42:25As mentioned, you guys will get
  1345. 42:26everything in the description if you
  1346. 42:28want it, and I'll also run you through
  1347. 42:30how to create it.
  1348. 42:31>> [gasps]
  1349. 42:31>> But for now, what I want to do, let's
  1350. 42:33say just hide this, is say something
  1351. 42:35along the lines of, "Hey,
  1352. 42:37build me a full-stack
  1353. 42:40app that lets users enter
  1354. 42:45a desired image to generate,
  1355. 42:48and then it generates said image. We'll
  1356. 42:50make this really simple because I don't
  1357. 42:52actually want this to take forever. I'm
  1358. 42:54kind of a time crunch today.
  1359. 42:56And I just want you guys to see how this
  1360. 42:58deals with that problem.
  1361. 43:00Keep in mind in this case, Claude, which
  1362. 43:03is the model that we're currently
  1363. 43:04talking to, cuz it's Claude Code, is
  1364. 43:06going to be our top-level orchestrator.
  1365. 43:10Okay?
  1366. 43:11Now, this is going to plan things out
  1367. 43:13for us, which is why it's entering this
  1368. 43:14plan mode.
  1369. 43:16Next, what we're going to do is we're
  1370. 43:17going to delegate all
  1371. 43:20difficult tasks, um, like back-end tasks
  1372. 43:23to Codex, as well as testing tasks.
  1373. 43:27Then down at the very bottom here, you
  1374. 43:28know, for anything related to front-end,
  1375. 43:31we're going to delegate that to Gemini.
  1376. 43:34And so we're going to build basically an
  1377. 43:35ecosystem here where Claude is shuttling
  1378. 43:37information back and forth between, uh,
  1379. 43:40you know, Codex and Gemini for various
  1380. 43:42things. And as you can see here, it's
  1381. 43:43already starting to ask me, "Hey, which
  1382. 43:45image generation API would you like to
  1383. 43:46use?" I'm actually just going to say,
  1384. 43:48um, Nano Banana Pro 2. It's a Google
  1385. 43:52product.
  1386. 43:54Okay, I'm going to submit that.
  1387. 43:55And now what it's going to do is it's
  1388. 43:57going to decide, "Hey, how am I going to
  1389. 43:59delegate this work?" At the end of it,
  1390. 44:01Claude will give me a plan, and you can
  1391. 44:02see here that it's decided on back-end,
  1392. 44:04front-end, and so on and so forth. And
  1393. 44:06what it'll do now is it'll actually
  1394. 44:08dispatch work to Gemini, Codex, and then
  1395. 44:11itself to fix a various integration
  1396. 44:13issues. So, I'm just going to say plan
  1397. 44:14approved, and now it's going to start
  1398. 44:16doing the coding. The way that Claude
  1399. 44:17Code does this is it uses the execute
  1400. 44:20task path for Codex. And so, what is
  1401. 44:23occurring right now is it's just sent
  1402. 44:25this big request in to Codex's best
  1403. 44:27model. Okay, and now just clicking the
  1404. 44:28button in the top right-hand corner, we
  1405. 44:30now have a preview. And um in this case,
  1406. 44:32Claude is now reviewing the generated
  1407. 44:34application and doing some self-testing.
  1408. 44:36And so, we built this image generator
  1409. 44:38app. We've asked for a cute cat wearing
  1410. 44:41sunglasses on a beach. This is now
  1411. 44:43passing through to an API that Claude
  1412. 44:46Code set up with a Gemini for the
  1413. 44:49front-end and then Codex for the
  1414. 44:50back-end's help. It's actually doing the
  1415. 44:52the generation right now. And we've
  1416. 44:53generated the cute picture of the cat on
  1417. 44:55the beach. Looks great to me.
  1418. 44:57The reason why you might want to do this
  1419. 44:58is because well, it's kind of twofold.
  1420. 45:00One, you get to parallelize your work as
  1421. 45:02mentioned. And so, you get to build the
  1422. 45:03front-end um using a model for which the
  1423. 45:05front-end builder is the best. You get
  1424. 45:07to build a back-end simultaneously using
  1425. 45:09model by which the back-end builder is
  1426. 45:11the best. And then you get to use an
  1427. 45:12orchestrator, which basically eeks out a
  1428. 45:14few percentage points increased like
  1429. 45:16reasoning and decision-making and stuff
  1430. 45:17like that because
  1431. 45:19it's able to evaluate the code from both
  1432. 45:21of these things independently without
  1433. 45:23being polluted by the context window.
  1434. 45:24And we're going to talk more about that
  1435. 45:25specific review pattern later. But um
  1436. 45:27this allows you to eke out, you know,
  1437. 45:28more quality. The downside of this um
  1438. 45:30prompt approach is it usually costs more
  1439. 45:33because now you're splitting your tokens
  1440. 45:34across multiple models just one
  1441. 45:36provider. And usually providers will
  1442. 45:37subsidize your token usage like Claude
  1443. 45:40will subsidize most of its usage on the
  1444. 45:41max plan for instance.
  1445. 45:43Um the $200 a month that you spend on it
  1446. 45:45is actually equivalent to like $5,000 a
  1447. 45:47month in usage. Whereas when you build
  1448. 45:48via API, it's usually a little bit more
  1449. 45:50standardized. And then as a result of
  1450. 45:51that, you end up building way more. You
  1451. 45:53don't you don't get that cool
  1452. 45:54subsidization. However, this is
  1453. 45:56something that people are increasingly
  1454. 45:57using for more complicated
  1455. 45:58infrastructural projects, especially
  1456. 46:00when as mentioned a minor percentage
  1457. 46:02point or two difference in terms of
  1458. 46:04quality is very important to you. And
  1459. 46:06so, this is me just doing this in
  1460. 46:07Claude, but you can obviously use, I
  1461. 46:08don't know, Codex as the orchestrator if
  1462. 46:10you wanted to build this in Codex. You
  1463. 46:11could use Gemini as the orchestrator if
  1464. 46:13you wanted to do this in, you know,
  1465. 46:14entirely Gemini. Right now, this is the
  1466. 46:16stack that seems to make the most sense,
  1467. 46:18what people are talking about the most.
  1468. 46:19If you guys are interested, the way that
  1469. 46:20all of this stuff works under the hood
  1470. 46:22is we basically set up a bunch of
  1471. 46:24different servers that call Codex and
  1472. 46:27Gemini inside of Claude. And so, that's
  1473. 46:29why we see this using the Claude
  1474. 46:31formatting above. It's because that
  1475. 46:32Claude is the orchestrator that's sort
  1476. 46:34of setting it up initially. And there's
  1477. 46:35also a Claude.md, which describes how
  1478. 46:37it's the manager. You know, you plan,
  1479. 46:39reason, delegate, validate, and fix
  1480. 46:41integration issues. When you break tasks
  1481. 46:43down, break them into front and back end
  1482. 46:45and test subtasks, and then delegate
  1483. 46:47things as required. I'm going to include
  1484. 46:48this prompt as well as everything else
  1485. 46:50you need in order to do the same thing
  1486. 46:51I'm down below in the description. But
  1487. 46:53in order for this to work, you will, of
  1488. 46:54course, need API keys for various
  1489. 46:56platforms. And in order to get those,
  1490. 46:57you do have to sign up to typically
  1491. 46:58something a little bit different what we
  1492. 47:00signed up to before. And in order to
  1493. 47:01sign up to those, you do typically need
  1494. 47:03to go directly to the platform, create
  1495. 47:05an account, and then set up an API key.
  1496. 47:07So, you can see over here, that's what
  1497. 47:08I've done for Claude. And you can also
  1498. 47:10do the same thing for OpenAI and then
  1499. 47:12Gemini. Once you have those keys, you
  1500. 47:13would just give it to whatever model you
  1501. 47:14want to use to be the orchestrator, and
  1502. 47:16then it would set this whole thing up
  1503. 47:17for you, and then I'd be able to reason
  1504. 47:19and then communicate with different
  1505. 47:20models on your behalf. The next advanced
  1506. 47:22prompting technique is the video to
  1507. 47:24action pipeline. To make a long story
  1508. 47:26short, up until quite recently, AI
  1509. 47:28agents were forced to learn entirely
  1510. 47:30through text descriptions of stuff. And
  1511. 47:32the reason why is because multimodality,
  1512. 47:34like vision, usually, at least in the
  1513. 47:37context of video, was sort of out of
  1514. 47:39bounds. There was just no way that we
  1515. 47:41could feasibly take videos, which were
  1516. 47:43millions upon millions of tokens when
  1517. 47:45stitched together,
  1518. 47:46you know, into some text format that an
  1519. 47:48agent would understand.
  1520. 47:50Well, now agents can learn from the same
  1521. 47:51medium humans learn from. And we do so
  1522. 47:53by combining a little bit about what I
  1523. 47:55showed you guys earlier, okay?
  1524. 47:56Multi-agent MCP orchestration with this
  1525. 47:59idea of passing requests through the
  1526. 48:02Gemini API cuz Gemini has built-in
  1527. 48:05support for video now. Basically, uh you
  1528. 48:07know how videos are a certain number of
  1529. 48:09frames per second, like this video for
  1530. 48:11instance is 30 frames a second. You can
  1531. 48:13tell if you find a way to to slow it
  1532. 48:15down to like
  1533. 48:160.03.
  1534. 48:17I'll go literally one frame every 0.03
  1535. 48:20seconds or something like that. Well,
  1536. 48:21what this model does is it divides
  1537. 48:23videos into one frame per second
  1538. 48:25instead. It then analyzes the images in
  1539. 48:28succession and then uses a form of
  1540. 48:31descriptive prompting to break that down
  1541. 48:32into very, very clear steps.
  1542. 48:35So, basically what occurs is you'll feed
  1543. 48:36in something like a YouTube tutorial
  1544. 48:37URL. Claude will receive the URL but
  1545. 48:39cannot watch the video natively. So,
  1546. 48:41instead it'll call the Gemini API.
  1547. 48:44Gemini will watch the full video. Gemini
  1548. 48:46will then extract the step-by-step
  1549. 48:47instructions formatted as like a
  1550. 48:49numbered list that's hyper-precise and
  1551. 48:51hyper-specific.
  1552. 48:53The structured steps will return to
  1553. 48:54Claude via a very similar flow to what I
  1554. 48:57showed you guys with the design. And
  1555. 48:59then Claude will execute each using
  1556. 49:00hyper-specific tools. Maybe if you're
  1557. 49:02teaching somebody how to build something
  1558. 49:03on Blender or Figma or something like
  1559. 49:05that, you just give it access to the
  1560. 49:06toolkit and it does it. Then the final
  1561. 49:08result is the agent will have replicated
  1562. 49:10the tutorial end-to-end. And in that way
  1563. 49:12they can learn from the exact same
  1564. 49:13medium that that we learn. So, I'll show
  1565. 49:15you number one where I got inspiration
  1566. 49:17from this and then number two how to do
  1567. 49:19this for an actual task which in my case
  1568. 49:21is going to be building a simple flow
  1569. 49:22out in a no-code tool called N8N. So,
  1570. 49:24first the inspiration was Spencer
  1571. 49:26Sterling's post on X. He said he built
  1572. 49:29an agentic system that taught itself the
  1573. 49:31Blender donut tutorial by watching it on
  1574. 49:33YouTube. It watched the tutorials,
  1575. 49:35extracted the steps, filled in the gaps
  1576. 49:37in its own tooling, and completed the
  1577. 49:38entire thing autonomously.
  1578. 49:40And it's quite impressive to be honest.
  1579. 49:41Um anybody that's done any sort of 3D
  1580. 49:43design, myself included, will know that
  1581. 49:45like the uh way you learn how to build
  1582. 49:47things in Blender is you watch this one
  1583. 49:48specific tutorial that shows you how to
  1584. 49:50build a donut. And through this process
  1585. 49:52of building the donut, you learn about
  1586. 49:54like textures, you learn about various
  1587. 49:56shapes, you learn about how to modify
  1588. 49:58them and sculpt and paint and do all
  1589. 50:00this stuff. So, I made my own donut
  1590. 50:01personally a few years ago. I showed it
  1591. 50:04to all my friends, but I probably never
  1592. 50:05touch Blender again.
  1593. 50:07Well, the issue with knowledge like this
  1594. 50:09is it's obviously extraordinarily
  1595. 50:10visual, right? In order to really learn
  1596. 50:12something, you have to watch a video.
  1597. 50:13You can't really break all that down
  1598. 50:15into like hyper-specific text
  1599. 50:16instructions unless, you know, somebody
  1600. 50:18were to just like literally go
  1601. 50:20step-by-step. Step one, click this
  1602. 50:22button. Step two, rotate 0.283° to the
  1603. 50:25left. Step three, do this. So, there's a
  1604. 50:27fair amount of nuance and flexibility
  1605. 50:28there. And that's where video learning
  1606. 50:30comes in handy. Human beings learn
  1607. 50:32through video, obviously, but models
  1608. 50:33have a tough time doing it. And so, what
  1609. 50:35we do is we convert all of this into a
  1610. 50:37sequence of steps. We leave some steps a
  1611. 50:39little bit more vague, a little bit more
  1612. 50:40general, let the model have its own kind
  1613. 50:43of interpretability, and then give it
  1614. 50:44some way to like screenshot its results
  1615. 50:46to match it up to, you know, like the
  1616. 50:48frames in the video. And so, this fellow
  1617. 50:50here built this cool like workflow
  1618. 50:52building studio. It's sort of like his
  1619. 50:54own main operating system, I suppose.
  1620. 50:56That's what this is. It's not like an
  1621. 50:57app that he downloaded. It's something
  1622. 50:58that he built. And then he fed in this
  1623. 51:01along with the workflow I'm about to
  1624. 51:02show you to have it actually like build
  1625. 51:04the freaking thing. And it's
  1626. 51:05communicating with this app, Blender,
  1627. 51:07using what's called MCP, Model Context
  1628. 51:09Protocol, which is the same thing that
  1629. 51:10we use to communicate with the various
  1630. 51:12models like Gemini and the Codex
  1631. 51:14earlier. And you can get all that stuff
  1632. 51:15in the description down below as well.
  1633. 51:17So, I have this stored as a Claude skill
  1634. 51:20in video to action over here. So, if I
  1635. 51:23open this up and read the skill, you
  1636. 51:24could see here that it actually says,
  1637. 51:26"Extract actionable steps from YouTube
  1638. 51:28videos using Gemini video understanding.
  1639. 51:30Use when the user provides a YouTube
  1640. 51:31link it wants to learn procedures,
  1641. 51:33extract steps, understand visual
  1642. 51:34tutorials, or turn video content into
  1643. 51:36executable instructions." And so, what's
  1644. 51:38occurring is it'll basically take a
  1645. 51:39video, it'll download it for me, so then
  1646. 51:41I'll just be able to feed in a YouTube
  1647. 51:43URL, and then it'll convert that into
  1648. 51:45like a highly optimized series of steps
  1649. 51:47that, you know you would only really
  1650. 51:48know or be able to use through the
  1651. 51:50context of like an actual video. And so
  1652. 51:52to demonstrate what I've done here is
  1653. 51:54instead of using Gemini within
  1654. 51:56Antigravity, which is sort of the usual
  1655. 51:57design pattern, I thought I'd show you
  1656. 51:59guys my actual stack like what I
  1657. 52:00personally use. I think it's much easier
  1658. 52:02if you just use the models inside of the
  1659. 52:04tools inside of the companies that made
  1660. 52:06them. But in my case I'm a very big fan
  1661. 52:08of this Antigravity
  1662. 52:10kind of container. Then inside of it I
  1663. 52:11use Claude code. And so in a way I'm
  1664. 52:13actually using a Google wrapper around a
  1665. 52:15Claude code or Anthropic extension and
  1666. 52:17that's communicating with a Claude or an
  1667. 52:19Anthropic model. If you guys want to
  1668. 52:21replicate the setup is as simple as just
  1669. 52:23opening up Antigravity, heading to the
  1670. 52:24left-hand side where it says extensions,
  1671. 52:27downloading the Claude code for VS code
  1672. 52:29plugin. I know it says VS code, don't be
  1673. 52:30confused this is very similar to
  1674. 52:32Antigravity, installing it and then you
  1675. 52:35also have to log in here. After you're
  1676. 52:37done you will have the exact same
  1677. 52:38functionality that you have in the
  1678. 52:39Claude desktop app that I just showed
  1679. 52:41you guys earlier when we built out that
  1680. 52:42little full stack app. Uh it's just
  1681. 52:44you'll have it within Antigravity which
  1682. 52:45also allows you to do things like you
  1683. 52:47know organize your files and stuff on
  1684. 52:48the left-hand side. So that's my
  1685. 52:50personal stack. You don't have to use
  1686. 52:51it. Some people judge me for it.
  1687. 52:53Whatever, I like it, it works for me.
  1688. 52:55Okay, so what I'm going to do is I'm
  1689. 52:57going to find a YouTube video that I
  1690. 52:58like and then I'm just going to feed it
  1691. 52:59in these instructions. So I'll say I
  1692. 53:00want you to use the video to action
  1693. 53:03pipeline on and then I'm going to go
  1694. 53:05grab an image. Now what I've done is
  1695. 53:06I've found a flow that I built forever
  1696. 53:08ago. It's a short video about 21 minutes
  1697. 53:10that shows you how to scrape leads
  1698. 53:11without paying for a few APIs. I'm going
  1699. 53:13to bring that back into my Antigravity
  1700. 53:15instance and then I'm going to do this.
  1701. 53:18And what this is going to do is it'll
  1702. 53:19start by invoking the skill and this is
  1703. 53:21the UX for skill invocation. I think
  1704. 53:23that's what it's called in English. Holy
  1705. 53:25crap, that better be what it's called in
  1706. 53:26English. And then it's now going to send
  1707. 53:29that over to Gemini then receive back a
  1708. 53:32list of highly specific instructions
  1709. 53:34that you know understand UX I don't know
  1710. 53:36highlight the colors of buttons and
  1711. 53:38stuff like that and so on and so forth
  1712. 53:39before actually running it locally on my
  1713. 53:41computer. At the end of it, you'll get a
  1714. 53:43super in-depth analysis that looks like
  1715. 53:45this. So, you can actually see down over
  1716. 53:47here, it says, "Here's the hyper
  1717. 53:49detailed breakdown with literally every
  1718. 53:51single step." I mean, like, "Hey,
  1719. 53:52navigate over to this thing at 17
  1720. 53:55seconds. Here's how to do this thing on
  1721. 53:56that." And and so on and so on. So,
  1722. 53:59like, it'll it'll literally it'll go
  1723. 54:00visually as well and actually tell us
  1724. 54:02what the end-to-end flow is going to
  1725. 54:03look like, but then we'll also have just
  1726. 54:05a tremendous amount of context about
  1727. 54:06everything. Um so, what we're going to
  1728. 54:08do now is we're going to feed that in
  1729. 54:09and actually have this control my
  1730. 54:10browser. So, I'm going to open up a new
  1731. 54:12Claude code instance by clicking that
  1732. 54:13little button above. We'll go bypass
  1733. 54:15permissions. Then I'll say, "Use G Maps
  1734. 54:17Scraper deep analysis.md
  1735. 54:21to build out the same end-to-end flow
  1736. 54:24for me." It's now going to open up a
  1737. 54:26Chrome DevTools MCP server. It's then
  1738. 54:28going to link that up to the end-to-end
  1739. 54:30account. Now, it's actually thinking
  1740. 54:31through everything that it's going to do
  1741. 54:33using this file as a reference. And now
  1742. 54:35it'll go through and actually control my
  1743. 54:36browser to do the build. For simplicity,
  1744. 54:38I'm just going to move this over to the
  1745. 54:40right. Okay, and as we see, it just laid
  1746. 54:42out the entire thing from left to right.
  1747. 54:45So, it went through. It then identified
  1748. 54:48what all of the steps were. It then
  1749. 54:49created it inside of its own little
  1750. 54:51conversation thread. And then it
  1751. 54:54essentially generated what's called
  1752. 54:55workflow JSON and then pasted it in.
  1753. 54:57Now, this can obviously interact with my
  1754. 54:58my browser as well. That's what it just
  1755. 55:00did. So, it just went to the top and
  1756. 55:01then basically imported this. What it's
  1757. 55:03going to do now is just make some finer
  1758. 55:05final minor changes. I'm going to
  1759. 55:07configure the Google Sheets node and
  1760. 55:08then we'll be on our way. So, what I'll
  1761. 55:09do is I'll just take a screenshot of
  1762. 55:11this and then paste it in. Then I'll
  1763. 55:12say, "You're connected." Now, it's just
  1764. 55:14going through and then it's selecting
  1765. 55:15various elements. So, in this case, it's
  1766. 55:17selecting that little search button.
  1767. 55:18It's uh mapping the the fields and stuff
  1768. 55:21like that. And then it'll just continue
  1769. 55:22testing this nonstop until I have a
  1770. 55:24working flow. You could see, you know,
  1771. 55:25just kind of I mean, I should be moving
  1772. 55:27this around cuz it's going to get
  1773. 55:28confused. But you could see that it's um
  1774. 55:30actually gone through and then pumped in
  1775. 55:32like a specific search term. It's It's
  1776. 55:34gone gone and basically done everything
  1777. 55:36for me. Really, the only thing left is
  1778. 55:37to do some sort of testing. You can see
  1779. 55:38that uh if we actually click execute
  1780. 55:40workflow, I'm just going to stop it here
  1781. 55:41so I don't consume anything else. It's
  1782. 55:43actually gone through and literally like
  1783. 55:44scraped Google Maps for us, which is
  1784. 55:46sweet. And it's just done so entirely by
  1785. 55:48watching the video. So, it's entirely
  1786. 55:49like native video understanding. And
  1787. 55:51then it's extraordinarily detailed
  1788. 55:53because we're we're dumping it all into
  1789. 55:55a file, and then it can just constantly
  1790. 55:56reference that file.
  1791. 55:58It's then doing kind of a combination of
  1792. 55:59like, I don't know, like ASCII or or
  1793. 56:02text-based markup to
  1794. 56:04uh you know, understand both the
  1795. 56:05structure at like a micro level and then
  1796. 56:06also like a macro level. Next, I want to
  1797. 56:08chat this idea of stochastic multi-agent
  1798. 56:11consensus.
  1799. 56:12In case you guys didn't know, if you
  1800. 56:14were to take one model, let's say Gemini
  1801. 56:173.1 Pro High, and if you were to ask it
  1802. 56:19like an idea question, "A, give me 10
  1803. 56:22ideas to do X, Y, and Z."
  1804. 56:25Every time you ask Gemini 3.1 Pro the
  1805. 56:29same thing, it'll return a slightly
  1806. 56:31different answer.
  1807. 56:32Now, this property, some call it
  1808. 56:34randomness, but I think the correct
  1809. 56:36technical term is stochasticity,
  1810. 56:38which is just where, due to minor
  1811. 56:39statistical variations in the input or
  1812. 56:42in the way that the models work, the
  1813. 56:44output is going to be slightly different
  1814. 56:46every time.
  1815. 56:47The reason why this is so valuable is
  1816. 56:49because you can exploit this tendency to
  1817. 56:51get much, much better answers.
  1818. 56:54For instance, let's say I run three
  1819. 56:56times. One,
  1820. 56:57two, and three.
  1821. 57:00The reality is, if I run a query that at
  1822. 57:02the very beginning says, "Give me three
  1823. 57:05ideas for X." Okay?
  1824. 57:08On the very first time, okay, we might
  1825. 57:11get idea A,
  1826. 57:13idea B,
  1827. 57:14and idea C.
  1828. 57:16If we were to hypothetically run this
  1829. 57:17again, we'd probably get idea A,
  1830. 57:20idea B,
  1831. 57:21but just due to statistical variation,
  1832. 57:24there is a chance that on the second
  1833. 57:26run, it won't deliver us idea C at all.
  1834. 57:28It'll actually deliver us idea D.
  1835. 57:30And on the third run, maybe we do B,
  1836. 57:33Maybe we do C, and then maybe we also do
  1837. 57:35E.
  1838. 57:36What stochastic multi-agent consensus
  1839. 57:38is, you basically automate the process
  1840. 57:40of spawning multiple agents, giving them
  1841. 57:43slightly varied input prompts to take
  1842. 57:45advantage of stochasticity, and then
  1843. 57:46instead of just getting, let's say,
  1844. 57:48three ideas, A, B, and C,
  1845. 57:50you get to exploit stats to get all of
  1846. 57:53the possibilities, including ones that
  1847. 57:55might be a little rarer the model is
  1848. 57:57less likely to actually answer with.
  1849. 57:59And so in this way, you get A, you get
  1850. 58:01B, you can get C, but you can also get
  1851. 58:03D, and then you can get E. And so, you
  1852. 58:05know, if you compare it to just one
  1853. 58:06naive search, what we've done is we
  1854. 58:08basically almost doubled the scope of
  1855. 58:11the ideation. Now, mathematically, this
  1856. 58:13is termed traversing the search space. I
  1857. 58:15want you to pretend hypothetically that
  1858. 58:17this like little pie chart here
  1859. 58:19represents all possible answers to a
  1860. 58:22question. Maybe the question is, I don't
  1861. 58:25know,
  1862. 58:26"What's the simplest way to get to 1
  1863. 58:27million subscribers?" Right? This is
  1864. 58:28something that I asked uh my my model a
  1865. 58:30little while ago, because I'm interested
  1866. 58:31in getting to 1 million subscribers.
  1867. 58:33Now, obviously, I'm not just doing what
  1868. 58:35the thing tells me, right? A lot of its
  1869. 58:36ideas are stupid. But if you think about
  1870. 58:38it, if I can parallelize a thousand
  1871. 58:39agents all coming up with their own
  1872. 58:41ideas, even if on net, the average reply
  1873. 58:43or idea is a little bit worse than
  1874. 58:45something I'd be able to do,
  1875. 58:47I still get to run it a thousand times,
  1876. 58:48right? It's like running like uh I don't
  1877. 58:50know, like a 90 Q
  1878. 58:52uh you know, it's like it's like
  1879. 58:52Einstein versus 10,000 95 IQ
  1880. 58:56researchers. It's like, well, the 10,000
  1881. 58:5895 IQ researchers, despite lacking the
  1882. 59:00brilliance of Einstein, they'll probably
  1883. 59:01statistically figure it out eventually,
  1884. 59:03right?
  1885. 59:04So, um if this whole pie chart, to go
  1886. 59:06back to things, is all possible
  1887. 59:08responses, if you just run one search,
  1888. 59:10basically what you're doing
  1889. 59:12is you're only actually getting like a
  1890. 59:14small chunk of all of the possibilities.
  1891. 59:16And so instead, what we're doing is
  1892. 59:17we're actually running multiple
  1893. 59:18searches, you know, one search is going
  1894. 59:20to get this, another search is going to
  1895. 59:21get that, another search is going to get
  1896. 59:23that, another search is going to get
  1897. 59:25that, and and and so on and so forth.
  1898. 59:27And then in this way, what we do next is
  1899. 59:29we take the answers and then the replies
  1900. 59:32of the model that should be red, and
  1901. 59:33this one should be blue.
  1902. 59:34And then in doing so, we get to traverse
  1903. 59:36significantly more of that search space
  1904. 59:38without actually necessarily consuming
  1905. 59:40any more of our time.
  1906. 59:41So, this is going to be kind of
  1907. 59:42difficult to understand, and I think
  1908. 59:44I've run out of colors here uh unless
  1909. 59:45you've done something like this before,
  1910. 59:48but I'll make it really simple by
  1911. 59:49actually giving you guys a brief
  1912. 59:50demonstration on, I don't know, some use
  1913. 59:52case or problem that uh I think we
  1914. 59:54probably all be able to relate to.
  1915. 59:56Another final benefit is you get to do
  1916. 59:58all this in parallel. So, like, you
  1917. 59:59know, if you think about it, if you were
  1918. 1:00:01to do one search and then do another
  1919. 1:00:03search afterwards, and then do another
  1920. 1:00:04search. So, for instance, let's say you
  1921. 1:00:06have a query, "Give me three ideas for
  1922. 1:00:07X." And then it gives you three ideas,
  1923. 1:00:10and you're like, "Yeah, I want another
  1924. 1:00:10three ideas." And it gives you another
  1925. 1:00:12three ideas, and you're like, "Yeah, I
  1926. 1:00:12want another three ideas." Well, at the
  1927. 1:00:14end of it, you may have, I don't know,
  1928. 1:00:15nine ideas or something, but it will
  1929. 1:00:16have taken a certain amount of time. If
  1930. 1:00:18the first search is 5 minutes, the
  1931. 1:00:19second search is 5 minutes, and the
  1932. 1:00:20third search is 5 minutes, well, you
  1933. 1:00:22just consumed 15 minutes, right? So,
  1934. 1:00:24instead, what this does is this just
  1935. 1:00:26copies the idea, okay? But then it
  1936. 1:00:29paralyzes it. So, "Hey, give me three
  1937. 1:00:31ideas for X." And then what we do is we
  1938. 1:00:32do one, two, and three, and in total
  1939. 1:00:35this takes 5 minutes. Then we just
  1940. 1:00:37combine those three answers back over
  1941. 1:00:38here. The formal way to do stochastic
  1942. 1:00:40multi-agent consensus, at least the way
  1943. 1:00:41that I'm doing it here, is we'll provide
  1944. 1:00:43a single question or prompt, then we'll
  1945. 1:00:45do slight framing variations of every
  1946. 1:00:47prompt that we're feeding into the
  1947. 1:00:48model, and then we'll feed in, I don't
  1948. 1:00:49know, I'll probably feed in like three
  1949. 1:00:51or four or five or maybe 10
  1950. 1:00:52simultaneously. Depends on how deep you
  1951. 1:00:54want it to go. And then um what will
  1952. 1:00:55happen is these will be instantiated
  1953. 1:00:57into what are called sub-agents, okay?
  1954. 1:00:59Which are similar to the main agent, but
  1955. 1:01:01they operate in their own defined
  1956. 1:01:02context window. And then all of these
  1957. 1:01:03will just report back their answers to
  1958. 1:01:05the parent agent. So, this parent over
  1959. 1:01:07here is basically going to
  1960. 1:01:09work with a whole fleet of sub-agents,
  1961. 1:01:12and then once they're all done their
  1962. 1:01:13work, it'll synthesize the answers. And
  1963. 1:01:15then because what we're looking for is
  1964. 1:01:16we're looking for like statistical
  1965. 1:01:17variation, it'll calculate um what's
  1966. 1:01:19called the mode, which is the frequency
  1967. 1:01:21of each answer, and then the median,
  1968. 1:01:22which is like the average of each
  1969. 1:01:23answer, before ultimately combining all
  1970. 1:01:25this to give you much better results.
  1971. 1:01:27One final idea there is this idea of
  1972. 1:01:28consensus. A lot of models are going to
  1973. 1:01:31say the same things, obviously. Some
  1974. 1:01:33models are going to say things that are
  1975. 1:01:34quite different. And then finally, there
  1976. 1:01:36will be outliers, which are wild cards.
  1977. 1:01:38These wild cards here potentially
  1978. 1:01:39brilliant, but they might only appear
  1979. 1:01:41like 5 or 10% of the time, which is why
  1980. 1:01:42we spawn so many of these agents that we
  1981. 1:01:44can actually like farm these wild cards.
  1982. 1:01:46We can we can milk them like cows. And
  1983. 1:01:48then in that way, you can up your best
  1984. 1:01:50ideas coming from these these fleets of
  1985. 1:01:51agents.
  1986. 1:01:53Um and then also save a lot of time in
  1987. 1:01:54things like product ideation. I don't
  1988. 1:01:56know, man, keyword search, titles for
  1989. 1:01:58for for content, at least that's what
  1990. 1:01:59I'm using it for, or a variety of other
  1991. 1:02:01things. How research inventions. I'm
  1992. 1:02:03sure Anthropic and and Google and OpenAI
  1993. 1:02:06probably have fleets of models that are
  1994. 1:02:07doing basically this exact same thing
  1995. 1:02:09behind the scenes constantly. So, let me
  1996. 1:02:10actually show you guys what this looks
  1997. 1:02:11like in practice. I'm just going to zoom
  1998. 1:02:13way out of this and close a bunch of
  1999. 1:02:15these so you don't have to look at them
  2000. 1:02:16anymore. Then I'm going to spawn a new
  2001. 1:02:18Claude code tab over here on the right.
  2002. 1:02:20And what I'm going to do is I'm going to
  2003. 1:02:21use this skill that I've set up called
  2004. 1:02:22stochastic multi-agent consensus. So,
  2005. 1:02:24opening this up so you guys can read it.
  2006. 1:02:26What we're doing is responding N agents,
  2007. 1:02:29where N is just the number that you
  2008. 1:02:30specify, with slight framing variations
  2009. 1:02:33to independently analyze a problem, then
  2010. 1:02:36aggregate results by consensus. We use
  2011. 1:02:39this for decision making, ranking
  2012. 1:02:40things, strategic analysis, or any
  2013. 1:02:43problems where you want to filter
  2014. 1:02:44hallucinations and surface high variance
  2015. 1:02:46ideas.
  2016. 1:02:48So, hypothetically, let's just say
  2017. 1:02:50"Hey, I've struggled a lot with finding
  2018. 1:02:52any traction on TikTok whatsoever. I've
  2019. 1:02:54built up a bunch of accounts, and I
  2020. 1:02:56can't seem to get more than like 1,000
  2021. 1:02:58views per TikTok account. I'd like you
  2022. 1:03:00to use stochastic multi-agent consensus
  2023. 1:03:02to help me come up with possible
  2024. 1:03:04candidate ideas to solve this. I'm going
  2025. 1:03:06to feed this idea in, okay?" And this is
  2026. 1:03:08a real idea, actually. We are struggling
  2027. 1:03:10to get
  2028. 1:03:11traction on TikTok. For whatever reason,
  2029. 1:03:13we got 450k followers on Instagram, no
  2030. 1:03:16problem, but, you know, the second we
  2031. 1:03:18move things over to TikTok, we're just
  2032. 1:03:19not really getting too many views.
  2033. 1:03:21So, what it's going to start with is it
  2034. 1:03:22will spawn 10 agents, all independently
  2035. 1:03:25analyzing my TikTok problem. And every
  2036. 1:03:27one of them will get slightly different
  2037. 1:03:29analytical framing to maximize the
  2038. 1:03:31diversity of ideas.
  2039. 1:03:33Just going to zoom in here so you guys
  2040. 1:03:34could see this, but we now have a
  2041. 1:03:36conservative analysis. So, Nick Suriya
  2042. 1:03:38has 287K YouTube subscribers. You know,
  2043. 1:03:42his YouTube audience is primarily
  2044. 1:03:43professionals. Here's a bunch of
  2045. 1:03:45information about him. He has a small
  2046. 1:03:47team. Here's how he's doing things and
  2047. 1:03:49and so on and so forth. This agent over
  2048. 1:03:51here says, "Hey, I want you to assume
  2049. 1:03:53limited time and budget." This agent
  2050. 1:03:55over here, "I want you to only focus on
  2051. 1:03:57what is measurable and provable."
  2052. 1:03:59This agent over here, you know, I want
  2053. 1:04:01you to think about it from the end user
  2054. 1:04:02and viewer perspective.
  2055. 1:04:04And so, what we're doing is we're
  2056. 1:04:04basically taking advantage of the
  2057. 1:04:05parallelizability of models, not
  2058. 1:04:08necessarily the base intelligence,
  2059. 1:04:09though the intelligence is obviously
  2060. 1:04:10important, but like we care more about
  2061. 1:04:12like scanning and searching through
  2062. 1:04:13space of all possible solutions really
  2063. 1:04:15quickly. And then at the end, we're
  2064. 1:04:17going to converge all this back with our
  2065. 1:04:18parent agent. Now, once all these agents
  2066. 1:04:20have turned green here, if I open up
  2067. 1:04:22this thinking tab, you could see that
  2068. 1:04:23it's now combining all of the
  2069. 1:04:25information from each individual one.
  2070. 1:04:28So, there's a bunch of suggestions
  2071. 1:04:29saying, "Hey, you should try fresh
  2072. 1:04:30account. You should try a device reset.
  2073. 1:04:32You should try clean fingerprinting.
  2074. 1:04:34Hey, you should try TikTok native hook
  2075. 1:04:36reformatting. Hey, you should do duets
  2076. 1:04:37with existing creators. Take advantage
  2077. 1:04:39of the fact that you're probably bigger.
  2078. 1:04:41Hey, you should do a series format, high
  2079. 1:04:42posting frequency, and so on and so
  2080. 1:04:44forth." Then you have some disagreements
  2081. 1:04:46here as well. And these disagreements
  2082. 1:04:47might be paid TikTok spark ads. Only one
  2083. 1:04:49of the 10 agents suggested something.
  2084. 1:04:51You know, in this one, they recommend
  2085. 1:04:53using shorts, but then in this one, they
  2086. 1:04:55recommend using a micro topic focus to
  2087. 1:04:57build authority and audience clarity.
  2088. 1:04:59You know, I'm not going to sit here and
  2089. 1:05:00pretend like all these ideas are the
  2090. 1:05:01bee's knees. Not all of them are
  2091. 1:05:03capturing lightning in a bottle, but you
  2092. 1:05:05run this thing long enough and you'll
  2093. 1:05:06see eventually you will get some pretty
  2094. 1:05:08good ideas. And the ideas will be
  2095. 1:05:10consensus ideas, like the idea of a
  2096. 1:05:11fresh account, but it'll also be kind of
  2097. 1:05:13like outlier ideas with pain point
  2098. 1:05:15framing, paid TikTok spark ads, niching
  2099. 1:05:17down your account identity,
  2100. 1:05:18cross-posting your Instagram Reels to
  2101. 1:05:20YouTube Shorts first. I mean, there
  2102. 1:05:21there there are a lot of possible ideas,
  2103. 1:05:23right?
  2104. 1:05:24>> [gasps]
  2105. 1:05:24>> Now, it's opened up this consensus
  2106. 1:05:26report, which I can visualize for you
  2107. 1:05:28guys by clicking this button.
  2108. 1:05:30And you can see here it's now saying,
  2109. 1:05:31"Hey, here is the context. TikTok growth
  2110. 1:05:33stalled at 1K views per account across
  2111. 1:05:35multiple accounts despite this massive
  2112. 1:05:37YouTube subs and 450,000 followers with
  2113. 1:05:39almost 5 million Reels views a month."
  2114. 1:05:41And then here,
  2115. 1:05:43this orchestrator now summarizes it and
  2116. 1:05:45says, "Hey, every agent independently
  2117. 1:05:47identified TikTok native hook
  2118. 1:05:48reformatting is really critical."
  2119. 1:05:50You know, Instagram is a little bit
  2120. 1:05:52different from TikTok hooks. Content
  2121. 1:05:54optimized for Instagram will
  2122. 1:05:55systematically fail TikTok's cold start
  2123. 1:05:57test. So, you actually have to
  2124. 1:05:58restructure it if you really want to
  2125. 1:05:59crush. Same thing here, fresh account,
  2126. 1:06:01clean device fingerprint. I mean, there
  2127. 1:06:03is just so much context here, it's not
  2128. 1:06:05even funny.
  2129. 1:06:06And so, the reality is I would have come
  2130. 1:06:07up with these ideas at some point, but I
  2131. 1:06:10basically got to put, you know, a genie
  2132. 1:06:12in a bottle and then have 500 genies
  2133. 1:06:16simultaneously solve my wishes at 100x
  2134. 1:06:19speed, and then aggregate all results
  2135. 1:06:21for um, you know, I don't know, probably
  2136. 1:06:23like three or four dollars realistically
  2137. 1:06:25in terms of tokens.
  2138. 1:06:27You also had a couple agents that said,
  2139. 1:06:28"Is TikTok even worth it?" And uh, I
  2140. 1:06:31think that's a really good question to
  2141. 1:06:32ask because up until now, I really
  2142. 1:06:34didn't think it was worth it. And so, in
  2143. 1:06:35general, anytime that I recommend you
  2144. 1:06:37have a strategic decision that you need,
  2145. 1:06:40you can make a quick one-time trade-off
  2146. 1:06:42of money for analysis by spawning a
  2147. 1:06:45bunch of agents all with slight prompt
  2148. 1:06:47variations,
  2149. 1:06:48and then collecting the rankings
  2150. 1:06:50reasoning to build this consensus map
  2151. 1:06:52document. And from here, you can figure
  2152. 1:06:54out your consensus items, your divergent
  2153. 1:06:56items, and then your outliers.
  2154. 1:06:59And you know, if they're consensus
  2155. 1:07:00items, well, odds are probably because a
  2156. 1:07:02lot of models have thought it's a good
  2157. 1:07:03idea, you should probably do it. If
  2158. 1:07:05there's some divergent items, well, you
  2159. 1:07:06should probably like reason about these
  2160. 1:07:08quite a bit before deciding whether it
  2161. 1:07:10makes sense. And if it's like an outlier
  2162. 1:07:12item, if there's only one out of 10
  2163. 1:07:13agents doing it, well, it can either be
  2164. 1:07:15a brilliant idea, in which case maybe
  2165. 1:07:17you should give it a try, or it might
  2166. 1:07:18just be a hallucination or some BS, in
  2167. 1:07:20which case you don't. And so, what this
  2168. 1:07:22allows you to do is execute with high
  2169. 1:07:23confidence. Thank you very much, AI, for
  2170. 1:07:25drawing that cute little That is a huge
  2171. 1:07:28fist. That thing would be terrifying in
  2172. 1:07:29real life. Um you know, this lets you
  2173. 1:07:32scan a large portion of the search space
  2174. 1:07:33in a very short period of time.
  2175. 1:07:35And uh yeah, the actual way that you
  2176. 1:07:37build it is very straightforward, and
  2177. 1:07:38I'll run you guys through what all that
  2178. 1:07:39stuff looks like I'm down below in the
  2179. 1:07:41project description. So, just like
  2180. 1:07:42stochastic multi-agent consensus allowed
  2181. 1:07:45us to scan large amounts of search space
  2182. 1:07:47in a short period of time. What we did
  2183. 1:07:49is we independently delegated work over
  2184. 1:07:51to agents and had them uh do things for
  2185. 1:07:53us. So too can we take advantage of the
  2186. 1:07:56same idea, but in my opinion get even
  2187. 1:07:58higher quality results through this idea
  2188. 1:08:00of agent chat rooms. What agent chat
  2189. 1:08:03rooms are are where instead of, you
  2190. 1:08:06know, parallelizing all the work and
  2191. 1:08:07having all these agents try and
  2192. 1:08:09independently solve problems, what you
  2193. 1:08:11do is you give all of them slightly
  2194. 1:08:12different personalities, and then you
  2195. 1:08:14have them all debate with each other
  2196. 1:08:16about these problems. And in doing so,
  2197. 1:08:18they tend to deliver much higher quality
  2198. 1:08:20responses because they're just like
  2199. 1:08:22they're they're a little bit spikier,
  2200. 1:08:23you know what I mean? They're not just
  2201. 1:08:24like a generalized idea, which I'll
  2202. 1:08:26visualize with like this interface, but
  2203. 1:08:28you know, because they're they're
  2204. 1:08:29butting heads with another, um
  2205. 1:08:31eventually they ideas get really nuanced
  2206. 1:08:33and really high quality. And so, um
  2207. 1:08:36whether or not you visualize things in
  2208. 1:08:37that way, that's personally how I think
  2209. 1:08:39about things. You really get to carve
  2210. 1:08:40out all the tiny little nooks and
  2211. 1:08:41crannies of an idea when you debate.
  2212. 1:08:44And so, here's a brief little
  2213. 1:08:45visualization. We start with a problem
  2214. 1:08:47or a prompt. We feed it in to, let's
  2215. 1:08:50say, three agents here, agent A, agent
  2216. 1:08:52B, and agent C.
  2217. 1:08:53All three are given the same document
  2218. 1:08:56called chat.json. And then what occurs
  2219. 1:08:58is they basically cycle through a debate
  2220. 1:09:00sequence where agent A says something,
  2221. 1:09:02agent B says something, and agent C says
  2222. 1:09:04something. And you know, if you do this
  2223. 1:09:06naively, the results will probably be
  2224. 1:09:07pretty low. But if you, I don't know,
  2225. 1:09:09force a little bit of a spark where
  2226. 1:09:11every agent has a slightly different
  2227. 1:09:12opinion and they're not afraid to like
  2228. 1:09:14state their opinion, um they'll
  2229. 1:09:15challenge each other's assumptions. They
  2230. 1:09:17will significantly improve the
  2231. 1:09:19probability that you catch errors. And
  2232. 1:09:21then this chat.json ends up being quite
  2233. 1:09:22a valuable resource because it also
  2234. 1:09:24shows like problem solving and stuff
  2235. 1:09:25like that. You can then give that to an
  2236. 1:09:27orchestrator and ultimately receive
  2237. 1:09:29higher quality output at the end. And so
  2238. 1:09:30it's sort of similar to what we had
  2239. 1:09:32earlier, right? It's just instead of
  2240. 1:09:33this operating um in parallel lanes,
  2241. 1:09:36what these agents are doing is actually
  2242. 1:09:38talking back and forth with each other.
  2243. 1:09:40And so they're actually capable of
  2244. 1:09:41having these conversations.
  2245. 1:09:42>> [sighs and gasps]
  2246. 1:09:43>> And I mean like I I just want you to
  2247. 1:09:44pretend we actually spawn 10 agents.
  2248. 1:09:46Agent one would be able to communicate
  2249. 1:09:47with agent two, but also agent three,
  2250. 1:09:49and also agent four, and also agent
  2251. 1:09:51five, and also agent six. So like the
  2252. 1:09:53total number of paths and um potential
  2253. 1:09:57like communication,
  2254. 1:09:58I don't really know what you want to
  2255. 1:09:59call them, like like vectors, um goes up
  2256. 1:10:02like crazy. And these agents,
  2257. 1:10:04ultimately, assuming that the idea is an
  2258. 1:10:06absolute BS, do end up at the end of it
  2259. 1:10:08like quite quite differentiated um in
  2260. 1:10:11their ideas and their opinions. So to
  2261. 1:10:13show you guys what this looks like, I
  2262. 1:10:14have another skill, which is just a
  2263. 1:10:16repeatable workflow, to be clear, where
  2264. 1:10:18I have this model chat. The description
  2265. 1:10:21here is to spawn five cloud instances on
  2266. 1:10:23a shared conversation room where they
  2267. 1:10:24debate, disagree, and converge on
  2268. 1:10:26solutions. They use round robin turns
  2269. 1:10:28with parallel execution within each
  2270. 1:10:30round for simplicity, and they trigger
  2271. 1:10:32on the model chat multi-model debate or
  2272. 1:10:34something else. So I have a bunch of
  2273. 1:10:36context down over here, and you guys can
  2274. 1:10:37grab this file for yourselves. What I'll
  2275. 1:10:39do is I'll actually just pipe this into
  2276. 1:10:40model chat. Okay, great. Use model chat
  2277. 1:10:44for a similar to really work through
  2278. 1:10:46this idea.
  2279. 1:10:47And now it'll spark this model chat
  2280. 1:10:50skill, which will then have them all
  2281. 1:10:52dump shared context into a little
  2282. 1:10:54chat.json, which I'll show you guys when
  2283. 1:10:56it's done. Okay, so the debate has now
  2284. 1:10:58concluded after these five agents had
  2285. 1:11:00this conversation. Okay, we can actually
  2286. 1:11:02see the the the chat conversation as
  2287. 1:11:05well by going down here to this model
  2288. 1:11:06chat. Uh, let's go latest and I'll go
  2289. 1:11:08conversation.
  2290. 1:11:09Um, basically what's occurred is we've
  2291. 1:11:11given it a topic to talk about and then
  2292. 1:11:14we've assigned a systems thinker, a
  2293. 1:11:15pragmatist, an edge case finder, a user
  2294. 1:11:17advocate, and then a contrarian to the
  2295. 1:11:19task. So, first of all, the systems
  2296. 1:11:21thinker begins, the pragmatist replies,
  2297. 1:11:23the edge case finder goes, the user
  2298. 1:11:24advocate goes, and so on and so forth.
  2299. 1:11:26And you can see each of them are um,
  2300. 1:11:28pretty pretty interestingly suggesting
  2301. 1:11:30uh, various approaches. So, these
  2302. 1:11:32advocates says, "Let me push back on
  2303. 1:11:33something that challenges the consensus
  2304. 1:11:35has glossed over, which is the clean
  2305. 1:11:36device plus fresh account fixes seems
  2306. 1:11:38fingerprinting is the problem. There's a
  2307. 1:11:40separate explanation nobody has stress
  2308. 1:11:41tested. Next content format is
  2309. 1:11:43fundamentally mismatched to TikTok's
  2310. 1:11:44cold start algo. And so, these are sort
  2311. 1:11:46of arriving at similar conclusions
  2312. 1:11:48despite the fact that uh, you know, we
  2313. 1:11:50instantiated this separately. And then
  2314. 1:11:52if we check out the synthesis, you can
  2315. 1:11:53see that all of them have agreed that we
  2316. 1:11:54need to run some diagnostics, that hook
  2317. 1:11:56reformatting is necessary but
  2318. 1:11:58sufficient. The high volume posting
  2319. 1:11:59blitz two to five a day is wrong, and
  2320. 1:12:01then fixing the IG YouTube pipeline
  2321. 1:12:03immediately is important regardless the
  2322. 1:12:05TikTok decision. This is something that
  2323. 1:12:07I guess it got context out from one of
  2324. 1:12:08my other files because um, basically
  2325. 1:12:10despite the fact that I have 450k
  2326. 1:12:12Instagram followers, very few of them
  2327. 1:12:13are converting to YouTube subscribers
  2328. 1:12:15and a lot of people, a lot of models as
  2329. 1:12:16well, are suggesting that the reason for
  2330. 1:12:18that is because Instagram is really
  2331. 1:12:19blocking outbound links, which I think
  2332. 1:12:22is actually fair. But then uh, there are
  2333. 1:12:23a lot of, you know, disagreements as
  2334. 1:12:25well. So, a lot of people say, "Nope,
  2335. 1:12:26stitch duet stupid. TikTok versus IG
  2336. 1:12:29pipeline is an either or. Device
  2337. 1:12:30fingerprinting might not be the issue,
  2338. 1:12:32maybe it's content mismatch, right?"
  2339. 1:12:34And uh, there are a lot of insights that
  2340. 1:12:36because we were able to sharpen our
  2341. 1:12:38opinions via debate,
  2342. 1:12:40these agents got that the previous model
  2343. 1:12:43runs through stochastic multi-agent
  2344. 1:12:45consensus did not. So, maybe we're
  2345. 1:12:47looking for saves, not completions.
  2346. 1:12:50Maybe there's just no category online
  2347. 1:12:51yet. And although this not true. If they
  2348. 1:12:53had the ability to research, they
  2349. 1:12:55probably would have figured this out.
  2350. 1:12:57Maybe it has to do with emotional
  2351. 1:12:59moments. And then here it even gave a
  2352. 1:13:00recommended execution plan.
  2353. 1:13:03So as mentioned, you know, I wouldn't
  2354. 1:13:04rely on agents for strategic advice at
  2355. 1:13:07the moment, but I would certainly not be
  2356. 1:13:09opposed to trading a little bit of my
  2357. 1:13:11money for a bunch of my time back and at
  2358. 1:13:13least ideating through the lower hanging
  2359. 1:13:15fruit. If you run enough of these
  2360. 1:13:17cycles, you will find pretty intriguing
  2361. 1:13:20and interesting outlier ideas. That's
  2362. 1:13:22just how statistics works. So you guys
  2363. 1:13:24can get all this down below in that
  2364. 1:13:25document. The next idea I want to talk
  2365. 1:13:27about is this idea of sub-agent
  2366. 1:13:29verification loops. To make a long story
  2367. 1:13:32short, where previously we took
  2368. 1:13:34advantage of parallelization, we're
  2369. 1:13:36going to take a step back now to sort of
  2370. 1:13:38serial processing.
  2371. 1:13:40But when an agent works really hard to
  2372. 1:13:43accomplish a task for you,
  2373. 1:13:45it usually gets pretty biased in that it
  2374. 1:13:49believes that its path was the best. And
  2375. 1:13:52the reason why is because, you know, it
  2376. 1:13:53just spent God knows how much time,
  2377. 1:13:55energy, and compute cycles building your
  2378. 1:13:58app or putting together your workflow or
  2379. 1:14:00doing your taxes or whatever the hell.
  2380. 1:14:03And because of that, you know, series of
  2381. 1:14:05like design decisions and then issues
  2382. 1:14:07and bug fixes, it's just very
  2383. 1:14:09consolidated in its opinion that the way
  2384. 1:14:11that it did what it did was the best. So
  2385. 1:14:14if you were to ask that same agent,
  2386. 1:14:15"Hey, can you make this better?" A lot
  2387. 1:14:17of the time it'll look at it and be
  2388. 1:14:18like, "Well, no, I did a pretty good
  2389. 1:14:20job. I don't think there's any way to do
  2390. 1:14:21it better."
  2391. 1:14:22However, instead of just giving that
  2392. 1:14:24agent back the entire context and
  2393. 1:14:26saying, "Can you do it better?" a much
  2394. 1:14:27smarter thing to do is to take all of
  2395. 1:14:29the outputs, not the reasoning, then
  2396. 1:14:32give the output, aka your code or your
  2397. 1:14:34workflow or the results of your your
  2398. 1:14:35accounting, to another agent and then
  2399. 1:14:38say, "Hey, is this right?" Because now
  2400. 1:14:40that second agent can evaluate purely
  2401. 1:14:42based off output. It doesn't actually
  2402. 1:14:44have to deal with evaluating things
  2403. 1:14:45based off the reasoning or the intent.
  2404. 1:14:47And so your work can end up being a lot
  2405. 1:14:50higher quality as a result.
  2406. 1:14:52So here's a quick example using like a
  2407. 1:14:53coding thing where we wanted to build a
  2408. 1:14:55rate limiter. What will happen is our
  2409. 1:14:57first agent will implement and write the
  2410. 1:15:00first draft of the code. This code
  2411. 1:15:02output will pass to a reviewer agent.
  2412. 1:15:05Now the reviewer agent is spawned with
  2413. 1:15:07fresh context, meaning there's no tokens
  2414. 1:15:09that are polluting its window. It has
  2415. 1:15:11zero bias. And what it does is just like
  2416. 1:15:13objectively speaking, you ask it, is
  2417. 1:15:15this thing correct? Are there any issues
  2418. 1:15:17here at first glance? Any ways you could
  2419. 1:15:19simplify this?
  2420. 1:15:21Now because it's treating this just like
  2421. 1:15:22it's treating random snippet of code it
  2422. 1:15:24finds on the internet, you know, it has
  2423. 1:15:26no opinions. It has no inherent like
  2424. 1:15:28desire to claim, well, this is the best
  2425. 1:15:31way because I spent all this time,
  2426. 1:15:32energy, and research figuring it out.
  2427. 1:15:34And it'll be able to to look at things
  2428. 1:15:35with, you know, those fresh eyes.
  2429. 1:15:37From there, if it finds issues, the idea
  2430. 1:15:39behind sub-agent verification loops is
  2431. 1:15:41it'll list those issues and then pass
  2432. 1:15:43the suggestions to a third agent called
  2433. 1:15:45a resolver, which has zero context about
  2434. 1:15:48any of this stuff as well. And so in
  2435. 1:15:49this way an implementer, reviewer,
  2436. 1:15:51resolver loop can get significantly
  2437. 1:15:54higher quality results than just one
  2438. 1:15:57agent doing everything simultaneously.
  2439. 1:15:59If there are no issues, everything's
  2440. 1:16:00approved, we're good to go.
  2441. 1:16:02Otherwise, it resolves, we do some
  2442. 1:16:04testing, and then we get the final
  2443. 1:16:05verified code output.
  2444. 1:16:07Are you guys noticing a trend here?
  2445. 1:16:09Basically, all of these like advanced
  2446. 1:16:11agent foundation advanced agent product
  2447. 1:16:13techniques ultimately circle back to
  2448. 1:16:16having multiple agents working in
  2449. 1:16:17parallel. And it's really interesting
  2450. 1:16:19because like the way that agents work
  2451. 1:16:21themselves is they already do work in
  2452. 1:16:22parallel. You know, a few years ago, um
  2453. 1:16:24agents were basically just one
  2454. 1:16:26statistical model, and you would ask the
  2455. 1:16:28statistical model to help you complete
  2456. 1:16:29the the the sentence or whatever, and
  2457. 1:16:31then it would give you the most likely
  2458. 1:16:32next token, and then it would rerun over
  2459. 1:16:34and over and over again until it did
  2460. 1:16:35that.
  2461. 1:16:36Well, a few years back, um people
  2462. 1:16:38started introducing this idea called a
  2463. 1:16:41mixture of experts, which is instead of
  2464. 1:16:43just having one model, what you do is
  2465. 1:16:45you actually send the same thing to like
  2466. 1:16:47three or four models, you average out
  2467. 1:16:50the statistical probabilities of every
  2468. 1:16:51word, and then you just pick what they
  2469. 1:16:53all converged on. Very similar to what I
  2470. 1:16:55did there with stochastic multi-agent
  2471. 1:16:56consensus. And so this mixture of
  2472. 1:16:58experts is sort of like the base
  2473. 1:17:00foundation that resulted in a really big
  2474. 1:17:02improvement in large language model
  2475. 1:17:04accuracy, among other things like
  2476. 1:17:06post-training and RLHF and and and stuff
  2477. 1:17:08like that. But what's really cool is all
  2478. 1:17:10of these frameworks basically do the
  2479. 1:17:12same idea. You know, we we treat these
  2480. 1:17:14mixture of experts now as themselves
  2481. 1:17:16models, and then we prompt them with
  2482. 1:17:19each other. We do them in parallel and
  2483. 1:17:21then integrate their answers like
  2484. 1:17:23stochastic multi-agent consensus. We
  2485. 1:17:24have them debate against each other like
  2486. 1:17:26with model chats. And now what we're
  2487. 1:17:28doing is we're basically having them
  2488. 1:17:29correct each other's work like with
  2489. 1:17:31sub-agent verification loops. So all of
  2490. 1:17:33these are just try
  2491. 1:17:34trading off the same core foundational
  2492. 1:17:37like features of models, which is that
  2493. 1:17:39at the end of the day they're
  2494. 1:17:39statistical machines. And so the more of
  2495. 1:17:41these statistics that you can, I don't
  2496. 1:17:43know, average out, the closer you get to
  2497. 1:17:45the reality. Another way of thinking
  2498. 1:17:46about this is if the implementer agent
  2499. 1:17:48has already spent 200,000 tokens
  2500. 1:17:50accumulating all that context, it'll
  2501. 1:17:52literally remember every wrong turn and
  2502. 1:17:54every dead end. It'll have a sunk cost
  2503. 1:17:56bias. It'll say, "Well, I wrote this, so
  2504. 1:17:57it must be right." And in a way it'll be
  2505. 1:17:59blind to its own mistakes. But you pass
  2506. 1:18:01it off to the super nerdy-looking
  2507. 1:18:03reviewer agent, it has a fresh empty
  2508. 1:18:05context. It'll only see the output, not
  2509. 1:18:07the journey that we took to get there.
  2510. 1:18:09No emotional attachment, although I
  2511. 1:18:10think this is unnecessary
  2512. 1:18:11anthropomorphization, and it'll catch
  2513. 1:18:13what the reviewer missed. So uh let me
  2514. 1:18:15show you guys how this actually looks
  2515. 1:18:17like in practice. Here I have this app
  2516. 1:18:19that I developed a while back for uh
  2517. 1:18:21video on vibe coding, and you guys can
  2518. 1:18:23check that out in the description if
  2519. 1:18:24you're interested. It's where I
  2520. 1:18:25basically put together a full end-to-end
  2521. 1:18:27system that allowed you to um design and
  2522. 1:18:29then syndicate a bunch of content. So,
  2523. 1:18:31you know, this is just some app, right?
  2524. 1:18:33This app, I don't even know if it's
  2525. 1:18:34fully functional. Okay, no, it isn't
  2526. 1:18:36because I had to turn it off. But,
  2527. 1:18:37hypothetically, there's a big code base
  2528. 1:18:39here, right? And so, what I want to do
  2529. 1:18:40is I want to use this app to show you
  2530. 1:18:42guys how an un
  2531. 1:18:44biased code reviewer would take a look
  2532. 1:18:46at the code that a previous agent had
  2533. 1:18:48written, in this case Gemini, and and
  2534. 1:18:50improve it. So, what I'm going to do is
  2535. 1:18:52I'm going to go find this repo. Okay,
  2536. 1:18:53and I found it over here. It's in the
  2537. 1:18:54Splinter repository. Uh makes sense. I'm
  2538. 1:18:57just going to open up a new Claude code
  2539. 1:18:58instance.
  2540. 1:18:59And then down over here, I'm going to
  2541. 1:19:00say, "I'd like you to use
  2542. 1:19:04And I just need to make sure I know what
  2543. 1:19:05the skill is called.
  2544. 1:19:08Agent review on the Splinter repo. It's
  2545. 1:19:12Let's just say folder. It's in the
  2546. 1:19:13parent folder. So, that it knows where
  2547. 1:19:15this is. Now, that way I can still
  2548. 1:19:17execute it within this business
  2549. 1:19:19workspace, which I found a much better
  2550. 1:19:21way of organizing things.
  2551. 1:19:22And while it's doing that, I'm going to
  2552. 1:19:23open up the skill.md.
  2553. 1:19:25So, what the skill.md does is it spawns
  2554. 1:19:28sub-agents to review, simplify, and
  2555. 1:19:30verify output. It uses after completing
  2556. 1:19:32any non-trivial implementation task, and
  2557. 1:19:34it triggers on the words review this,
  2558. 1:19:36agent review, self-review, or, you know,
  2559. 1:19:38{slash} agent-review. And you can see
  2560. 1:19:40it's already doing this. It's spun up a
  2561. 1:19:42sub-agent called review Splinter
  2562. 1:19:44codebase. And what this does is it
  2563. 1:19:46reviews it for four things: correctness,
  2564. 1:19:49edge cases, simplification, and then
  2565. 1:19:51security. Now, like, do I know how to do
  2566. 1:19:53all this programming under the hood? No,
  2567. 1:19:55I don't. But, these agents certainly do.
  2568. 1:19:57And so, we can take advantage of that by
  2569. 1:19:58having an agent with zero context, this
  2570. 1:20:00one here, review that entire workspace
  2571. 1:20:03sort of independently and objectively.
  2572. 1:20:05And now it's doing a bunch of reading,
  2573. 1:20:07and it's going to integrate that with
  2574. 1:20:08the suggestions of this model to give us
  2575. 1:20:10a much higher quality output. All right,
  2576. 1:20:12the Splinter code review just finished
  2577. 1:20:14up, and we found 22 issues across the
  2578. 1:20:16codebase. There's some critical ones
  2579. 1:20:17here, some high issues here, some medium
  2580. 1:20:20issues here, and then some low issues
  2581. 1:20:22over there. Now, it's asking me if I
  2582. 1:20:24want me to start fixing any of these,
  2583. 1:20:26and I'll say, "Absolutely." And the
  2584. 1:20:28whole idea behind this now is we're
  2585. 1:20:30we're capable of looking at this
  2586. 1:20:32completely objectively. You know, like I
  2587. 1:20:34asked the initial model Gemini when I
  2588. 1:20:36made the the app in the course like
  2589. 1:20:37multiple times, "Hey, are there any
  2590. 1:20:39issues here? Hey, are there any ways to
  2591. 1:20:40make this better? Hey, what do you
  2592. 1:20:41suspect is a problem?" And it just
  2593. 1:20:42couldn't find it because it was so
  2594. 1:20:43polluted by its own biases. Now, another
  2595. 1:20:46model can. And it's very similar to like
  2596. 1:20:49peer review in like academic um circles.
  2597. 1:20:52It's not that like, you know, you're
  2598. 1:20:53dumb for coming up with this code base.
  2599. 1:20:55Like, how dare you? It's just that as
  2600. 1:20:57you work on things more and more and
  2601. 1:20:59more, you tend to see things a little
  2602. 1:21:01more narrow and more narrow because
  2603. 1:21:02you've explored a bunch of other
  2604. 1:21:03possible paths. And the reality is
  2605. 1:21:06the fact that you explored those paths
  2606. 1:21:08and those don't work don't necessarily
  2607. 1:21:09mean that if somebody else explored one
  2608. 1:21:11of those paths, it wouldn't work either.
  2609. 1:21:13And so, this is just a way of remaining
  2610. 1:21:14as objective as humanly possible, which
  2611. 1:21:15is obviously a very valuable thing to do
  2612. 1:21:17when you're doing things like creating
  2613. 1:21:18applications, code, um you know, sales,
  2614. 1:21:21marketing, and all the various things
  2615. 1:21:22that AI agents allow us to do. Next up,
  2616. 1:21:24I want to talk a little bit about prompt
  2617. 1:21:26contracts. For those of you guys that
  2618. 1:21:28don't know, earlier on we chatted a
  2619. 1:21:29little bit about a definition of done,
  2620. 1:21:31right? Well, vague tasks, aka tasks that
  2621. 1:21:34don't have clearly defined definitions
  2622. 1:21:36of done, are basically the number one
  2623. 1:21:39problem nowadays with what I would
  2624. 1:21:42consider to be people's like disillusion
  2625. 1:21:43with AI agents. Like, when a total
  2626. 1:21:45novice starts using AI and then they
  2627. 1:21:47dive into some agent to coding platform
  2628. 1:21:49and then they just say, "Hey, build me a
  2629. 1:21:51Netflix 2.0. Make me a million dollars.
  2630. 1:21:54Make no mistakes."
  2631. 1:21:55Um because of their extraordinarily
  2632. 1:21:57poorly defined definition of done,
  2633. 1:21:59because of the poorly defined goals,
  2634. 1:22:01because they don't give it any
  2635. 1:22:02constraints, because they don't give it
  2636. 1:22:04any failure conditions,
  2637. 1:22:06uh that model is just not going to do
  2638. 1:22:07any any get anywhere near as high
  2639. 1:22:09quality end-end result as if they did
  2640. 1:22:11just follow a simple little uh
  2641. 1:22:12step-by-step process.
  2642. 1:22:14And so, the step-by-step process
  2643. 1:22:15obviously you could learn,
  2644. 1:22:17but you could also just like hard code
  2645. 1:22:19it as a skill somewhere in your
  2646. 1:22:20workspace or as uh you know, something
  2647. 1:22:22in your cloud and MD and they just force
  2648. 1:22:24your model to always have this
  2649. 1:22:25information before you proceed.
  2650. 1:22:27And so, for instance, if you give it a
  2651. 1:22:28vague task like build a rate limiter,
  2652. 1:22:30okay, it'll do pretty poorly. But, the
  2653. 1:22:32whole idea behind a prompt contract is
  2654. 1:22:34you basically make the user who puts in
  2655. 1:22:36a request like this sign a mini contract
  2656. 1:22:38and just say, "Okay, cool. The contract
  2657. 1:22:40is, you know, here's what your goal is.
  2658. 1:22:42Here what your constraints are. Here's
  2659. 1:22:44what your format is and here's what your
  2660. 1:22:45failure is. Are you good to go?" If the
  2661. 1:22:47answer to that question is yes, now the
  2662. 1:22:49model has actually gone through the step
  2663. 1:22:50of defining your goal, your constraints,
  2664. 1:22:52your format, and your failure. And so,
  2665. 1:22:55all of your definitions are done. All of
  2666. 1:22:57the various kind of technical spec
  2667. 1:22:59requirements here are much more laid
  2668. 1:23:01out. And then, the model sort of has a
  2669. 1:23:03lot easier of a way of going about
  2670. 1:23:04things.
  2671. 1:23:05And so, this is very similar, if you
  2672. 1:23:06guys are aware, to like this idea of
  2673. 1:23:09scopes.
  2674. 1:23:11Now, I, you know, I run like a freelance
  2675. 1:23:13education platform, like an AI
  2676. 1:23:15automation agency education platform.
  2677. 1:23:17And so, scopes are a really big part of
  2678. 1:23:18like a successful project.
  2679. 1:23:20And so, I teach people how to define
  2680. 1:23:21like really precise and concrete scopes,
  2681. 1:23:24whether you're doing, you know, a small
  2682. 1:23:25project for a client or working with
  2683. 1:23:27some large enterprise businesses or
  2684. 1:23:28something like that.
  2685. 1:23:29And like a real real common issue is
  2686. 1:23:31scopes just tend to either to be way too
  2687. 1:23:33vague.
  2688. 1:23:35And so, people don't actually clearly
  2689. 1:23:36define them.
  2690. 1:23:38Or, they end up way too restrictive. And
  2691. 1:23:41so far that people, you know, in a in an
  2692. 1:23:44attempt to counterbalance the vagueness,
  2693. 1:23:45they end up going like way too specific
  2694. 1:23:47and then the scope ends up being like so
  2695. 1:23:48restrictive that it's like, you know,
  2696. 1:23:50you're a slave to it and you can't
  2697. 1:23:51change anything.
  2698. 1:23:52And so, prompt contracts sort of help
  2699. 1:23:53you navigate the thin line between too
  2700. 1:23:56vague and too restrictive. And that's
  2701. 1:23:57very similar in nature to like giving a
  2702. 1:23:59contractor a task and then the
  2703. 1:24:01contractor clarifying with you before
  2704. 1:24:03they actually do the task, which I
  2705. 1:24:05think, you know, is clearly a
  2706. 1:24:08consequence of agents pushing all of us
  2707. 1:24:10more towards like management style
  2708. 1:24:11positions, where we just manage the
  2709. 1:24:13inputs and the outputs of these things.
  2710. 1:24:15So, I'm a big fan of defining these
  2711. 1:24:16clearly. So, what does this actually
  2712. 1:24:18mean in practice? Well, there's
  2713. 1:24:20obviously a million and one different
  2714. 1:24:21ways you can define prompt contracts.
  2715. 1:24:23The way that I've decided to do so in
  2716. 1:24:25this demonstration is through a skill
  2717. 1:24:27called prompt-contract.
  2718. 1:24:29And so basically before implementing any
  2719. 1:24:30non-trivial task, this skill forces you
  2720. 1:24:33to generate a structured prompt contract
  2721. 1:24:35with goals, constraints, the format of
  2722. 1:24:37output, and then failure.
  2723. 1:24:39So the idea here is you're treating it
  2724. 1:24:40just like a spec or a scope of work. Any
  2725. 1:24:43task that produces code or some
  2726. 1:24:45configuration settings or something like
  2727. 1:24:46that needs to go through this process.
  2728. 1:24:48And then this model will sort of
  2729. 1:24:49self-analyze the request before drafting
  2730. 1:24:52a four-section contract and then
  2731. 1:24:54presenting it for approval. This is
  2732. 1:24:56almost similar in nature to like the
  2733. 1:24:58plan mode that a lot of these agent
  2734. 1:25:00platforms now have. Like in Claude code
  2735. 1:25:02for instance, it can enter plan mode and
  2736. 1:25:03give you a brief little plan and have
  2737. 1:25:05you approve the plan before it proceeds.
  2738. 1:25:07It's just this formalizes it as a
  2739. 1:25:08contract. And no, you're not signing
  2740. 1:25:10your life away with Claude code when you
  2741. 1:25:11do this. But you know, it's a simple and
  2742. 1:25:13easy way to make sure that you get more
  2743. 1:25:15repeatable and consistent and accurate
  2744. 1:25:17outputs every time. So why don't I
  2745. 1:25:18actually do this? Use prompt contracts
  2746. 1:25:22to define this task. And then I'm just
  2747. 1:25:24going to pretend that I'm giving it a
  2748. 1:25:25really simple query. I'm just going to
  2749. 1:25:27say I want you to build me a beautiful
  2750. 1:25:30site for leftclick.ai. That's my agency.
  2751. 1:25:34So what it's going to do is it'll begin
  2752. 1:25:35by invoking the skill prompt contract.
  2753. 1:25:38And I mean beautiful site is such a
  2754. 1:25:39subjective term, right? I mean like what
  2755. 1:25:41the heck does that even mean? And so the
  2756. 1:25:42model is going to be essentially forced
  2757. 1:25:44to ask me for more context on what
  2758. 1:25:46constitutes a beautiful site to me. And
  2759. 1:25:49in this way I will get a much higher
  2760. 1:25:50quality site or app or whatever the hell
  2761. 1:25:52at the end of it. Likewise, you could do
  2762. 1:25:53this with any business task as well. It
  2763. 1:25:55doesn't just have to be like a design
  2764. 1:25:56task. You could set up a prompt contract
  2765. 1:25:59for hey, email these 45 people and it
  2766. 1:26:01could ask you like oh like what spec,
  2767. 1:26:02you know, specifications do you do you
  2768. 1:26:04want to confirm that they're emailed?
  2769. 1:26:06And uh what do you want the emails to
  2770. 1:26:07say? And what's the goal of a successful
  2771. 1:26:09thing? And like do you have any failure
  2772. 1:26:11parameters? If we only email 44, is that
  2773. 1:26:13okay with you? Right? It basically
  2774. 1:26:15forces it to be a lot more clear and
  2775. 1:26:16then concise.
  2776. 1:26:18So, what's happening now is it's gone
  2777. 1:26:19through and it's actually accessed
  2778. 1:26:20leftclick.ai. That's my current website,
  2779. 1:26:22and then it's getting a bunch of
  2780. 1:26:24screenshots and stuff like that. And the
  2781. 1:26:25reason why is because it's attempting to
  2782. 1:26:27build a context for the prompt contract.
  2783. 1:26:29So, its first step was to analyze the
  2784. 1:26:31request, right? What it's going to do is
  2785. 1:26:33it'll identify what Don looks like.
  2786. 1:26:35It'll identify some implicit
  2787. 1:26:36assumptions. So, what am I about to
  2788. 1:26:38force the model to assume without being
  2789. 1:26:39told? Well, obviously an assumption is I
  2790. 1:26:41already have a website, right? And so,
  2791. 1:26:42it's going to go through, take pictures
  2792. 1:26:44of my website, and see, well, if Nick
  2793. 1:26:46wants something different from this,
  2794. 1:26:47why? And then it's going to sort of make
  2795. 1:26:49its own judgment to that end. And now
  2796. 1:26:51it's actually giving me the contract.
  2797. 1:26:52So, the goal is a single-page marketing
  2798. 1:26:54site for Left Click. Here's some
  2799. 1:26:56constraints. You know, we want smooth
  2800. 1:26:58scroll animations under 500 lines of
  2801. 1:27:00HTML. The format is this. There should
  2802. 1:27:03be these sections. Subtle animations
  2803. 1:27:05fade in on scroll hover states. A
  2804. 1:27:07failure is if it looks like a generic
  2805. 1:27:09Bootstrap template. A failure is if it's
  2806. 1:27:11broken on mobile. A failure is if the
  2807. 1:27:13animations are janky. The failure is if
  2808. 1:27:15the file exceeds 500 lines. So, I
  2809. 1:27:16actually really like this prompt
  2810. 1:27:17contract. It's really simple and
  2811. 1:27:18straightforward. So, I'm actually going
  2812. 1:27:20to say go ahead and build it. But what's
  2813. 1:27:21cool is, you know, we're now actually
  2814. 1:27:22having a a conversation about this.
  2815. 1:27:24We're actually agreeing on, you know,
  2816. 1:27:26what the end result is going to be.
  2817. 1:27:28And this is actually really similar in
  2818. 1:27:30nature to the other thing that I want to
  2819. 1:27:31talk to you guys about, which is um kind
  2820. 1:27:34of related and orthogonal to prompt
  2821. 1:27:36contracts, although it is a little bit
  2822. 1:27:37different. And this is called a reverse
  2823. 1:27:39prompting. Now, reverse prompting is in
  2824. 1:27:42a similar vein, a mechanism used to
  2825. 1:27:44clarify the quality of a prompt and
  2826. 1:27:46improve the probability that it ends up
  2827. 1:27:48okay. And basically the way that this
  2828. 1:27:50works is instead of just like forcing
  2829. 1:27:51the model to give you this contract and
  2830. 1:27:53you having you sign off on it, it it
  2831. 1:27:55takes it one step further, actually
  2832. 1:27:56forces the model to ask you some
  2833. 1:27:58clarifying questions ahead of time. So,
  2834. 1:28:00rather than just give you a spec sheet
  2835. 1:28:01and say, "Okay, we're good to go." what
  2836. 1:28:02reverse prompting does is it has the
  2837. 1:28:05model ask you a bunch of questions that
  2838. 1:28:06you maybe didn't even think that you had
  2839. 1:28:08to answer. The model then takes all that
  2840. 1:28:10context and then feeds that into a
  2841. 1:28:11prompt contract later on. Okay, so step
  2842. 1:28:14one is when the user gives a task to an
  2843. 1:28:15AI agent. So I don't know, this is like
  2844. 1:28:17a website, right? Step two is the agent
  2845. 1:28:19asks five clarifying questions back to
  2846. 1:28:21the user before starting.
  2847. 1:28:23Step three is when we answer and then
  2848. 1:28:24the agent builds the correct thing on
  2849. 1:28:26the first try. So significantly improves
  2850. 1:28:27one-shot potential. And then if we
  2851. 1:28:29didn't have reverse prompting, there'd
  2852. 1:28:30be a lot of like wrong implicit
  2853. 1:28:32assumptions here, which would result in,
  2854. 1:28:33you know, the probability of a one-shot,
  2855. 1:28:36which is just when the agent does it in
  2856. 1:28:37literally one request,
  2857. 1:28:39going down quite a bit. And so
  2858. 1:28:41similarly, I also have a reverse prompt
  2859. 1:28:43skill over here. And so if I go to this
  2860. 1:28:46reverse prompt skill, you can see the
  2861. 1:28:47way that this is set up is before
  2862. 1:28:49implementing any non-trivial build, ask
  2863. 1:28:51the user five dynamically generated
  2864. 1:28:52clarifying questions to surface
  2865. 1:28:54non-obvious preferences, assumptions,
  2866. 1:28:55and constraints.
  2867. 1:28:57So when to trigger before starting
  2868. 1:28:58implementation, step one, analyze the
  2869. 1:29:00request, figure out some stated
  2870. 1:29:02requirements, implicit assumptions,
  2871. 1:29:05some decision points, failure modes, and
  2872. 1:29:06taste-dependent choices, right?
  2873. 1:29:08And so likewise, if I instead wanted to
  2874. 1:29:10build, let's say,
  2875. 1:29:12something for
  2876. 1:29:15build a beautiful site for One Second
  2877. 1:29:17Copy, which is my old content writing
  2878. 1:29:18company, which we just had to shut down
  2879. 1:29:19a few days ago.
  2880. 1:29:21Uh as you guys can imagine, content
  2881. 1:29:22isn't super in these days.
  2882. 1:29:25Oh, and then um use the reverse prompt
  2883. 1:29:28skill
  2884. 1:29:30and chain it together with prompt
  2885. 1:29:32contracts after.
  2886. 1:29:34What you could see is we're now engaged
  2887. 1:29:36significantly more than we were before.
  2888. 1:29:38Before, I just say, "Build me a
  2889. 1:29:40beautiful site." Probability that it
  2890. 1:29:41gets what I want right on the first try,
  2891. 1:29:43pretty damn low. What it's doing now is
  2892. 1:29:45it's asking a bunch of clarifying
  2893. 1:29:47questions to confirm whether or not, you
  2894. 1:29:49know, this site is as I want it to be.
  2895. 1:29:52And then after I feed it back that
  2896. 1:29:53information, it'll then take that and
  2897. 1:29:55use that to construct essentially that
  2898. 1:29:57that prompt contract that we had before.
  2899. 1:29:59So here's what the conversation looks
  2900. 1:30:00like. What's the primary goal of the
  2901. 1:30:02site? Brand credibility, sales funnel,
  2902. 1:30:04lead gen? You know, what I want is just
  2903. 1:30:05brand credibility. Should it be a single
  2904. 1:30:07static page site or should I build it in
  2905. 1:30:09some other framework? No, I wanted a
  2906. 1:30:10simple site. What's the vibe? You know,
  2907. 1:30:12it's AI content writing. You know,
  2908. 1:30:14should I do a clean modern SaaS
  2909. 1:30:16aesthetic like linear versus L? Do I
  2910. 1:30:18want something different? Yeah, I want
  2911. 1:30:20like linear but white.
  2912. 1:30:22You know, should I generate the copy
  2913. 1:30:23from context or use some placeholder
  2914. 1:30:25content? No, you're cool. You can
  2915. 1:30:26generate it from here.
  2916. 1:30:28Now, once we've clarified everything,
  2917. 1:30:30what this model is going to do is use
  2918. 1:30:32all this information to outline the
  2919. 1:30:33prompt contract using the prompt
  2920. 1:30:35contract skill. And now you can see it's
  2921. 1:30:36invoking the skill as well.
  2922. 1:30:38And here we have a contract. It'll be a
  2923. 1:30:40single page static site for one second
  2924. 1:30:41copy linear white aesthetic five
  2925. 1:30:43sections to play ready. Here's some
  2926. 1:30:45constraints.
  2927. 1:30:46Here's the format. Maybe I don't like
  2928. 1:30:47the format. Maybe I don't want it inside
  2929. 1:30:49of active. You know, I want it somewhere
  2930. 1:30:50else.
  2931. 1:30:51Uh but anyway, in in this case, maybe I
  2932. 1:30:53want to look good and and build it.
  2933. 1:30:55Now, just to show you guys an example of
  2934. 1:30:57how much higher quality we can get when
  2935. 1:30:58we actually do this. This is the um
  2936. 1:31:00website that uh it just built for us.
  2937. 1:31:02I'm just going to refresh this puppy and
  2938. 1:31:03take it to a new window because it gets
  2939. 1:31:05cut off on that window. This is here we
  2940. 1:31:07have those cool sexy animations. As we
  2941. 1:31:09scroll down, we also have some
  2942. 1:31:11information. Um it's it's light theme,
  2943. 1:31:13right? We have these really minimalistic
  2944. 1:31:15requirements here. Information about
  2945. 1:31:17myself, some services page, words from
  2946. 1:31:19happy clients, and then ultimately like
  2947. 1:31:20a CTA. And so, you know, the reason why
  2948. 1:31:23it was able to get much closer to what I
  2949. 1:31:24wanted, which was a minimalistic white
  2950. 1:31:26high-end aesthetic, is just because
  2951. 1:31:27like, you know, I I had it outlined in a
  2952. 1:31:29contract. As I'm sure you guys can
  2953. 1:31:30imagine, you can employ the same
  2954. 1:31:31approach for whatever the heck you want,
  2955. 1:31:33whether you're building a site or you
  2956. 1:31:34are you know, selling to people or you
  2957. 1:31:36are doing some sort of bookkeeping or
  2958. 1:31:38accounting. It's all just about uh
  2959. 1:31:39building out a very strong definition of
  2960. 1:31:41done. And the model can assist you with
  2961. 1:31:43this. You don't actually have to sit
  2962. 1:31:44down and laboriously write it all out
  2963. 1:31:45yourself. And that takes us to the
  2964. 1:31:47initial demo that we started with, which
  2965. 1:31:49was the multi-agent Chrome MCP manager.
  2966. 1:31:54Now, basically, at the very beginning of
  2967. 1:31:56this course, you didn't understand how,
  2968. 1:31:59you know, one agent could spawn a bunch
  2969. 1:32:00of other agents. You didn't understand a
  2970. 1:32:02lot of like the parallelization plays.
  2971. 1:32:04He also didn't understand that uh you
  2972. 1:32:06know, you could have agents actually
  2973. 1:32:08chat with each other and communicate. He
  2974. 1:32:10didn't understand the idea behind using
  2975. 1:32:12one agent to verify the work of another.
  2976. 1:32:14He didn't understand the idea behind
  2977. 1:32:15delegating to multiple different types
  2978. 1:32:17of models. What's really cool is the
  2979. 1:32:19multi-agent Chrome setup that I showed
  2980. 1:32:22you guys where we had, you know, five or
  2981. 1:32:2310 agents all operating independently in
  2982. 1:32:25their own browsers in their own
  2983. 1:32:26workspaces. All of that just feeds off
  2984. 1:32:29of this this idea or this concept um of
  2985. 1:32:32you know, agents increasing their level
  2986. 1:32:34of communication with other agents.
  2987. 1:32:36And so, essentially, if you think about
  2988. 1:32:38this logically, you know, if I were to
  2989. 1:32:39do this uh with like a single agent. So,
  2990. 1:32:41let's just say one agent.
  2991. 1:32:44You know, it's not actually rocket
  2992. 1:32:45science to have one agent use a browser
  2993. 1:32:47these days. There are built-in skills
  2994. 1:32:49called MCPs, model context protocols
  2995. 1:32:51basically, that you can just pipe in and
  2996. 1:32:53immediately connect to and it can do
  2997. 1:32:54everything for you, okay? It can It can
  2998. 1:32:56launch Chrome and then it can control
  2999. 1:32:57things on the page and and whatnot. It
  3000. 1:32:58can do that.
  3001. 1:33:00>> [gasps]
  3002. 1:33:00>> So, you know, the issue is it just takes
  3003. 1:33:02a lot of time. We'll receive the target
  3004. 1:33:04URL. We'll launch Chrome by the dev
  3005. 1:33:06tools MCP. We'll navigate to the
  3006. 1:33:07website. We'll take a screenshot. And
  3007. 1:33:10you know, in my case, this over here was
  3008. 1:33:11like um specific for me,
  3009. 1:33:14which was just page
  3010. 1:33:16or rather form fills.
  3011. 1:33:19After that, we'll identify the form,
  3012. 1:33:20extract the form fields, generate a
  3013. 1:33:21personalized message, fill the fields,
  3014. 1:33:23and then click submit. Um but you know,
  3015. 1:33:25this is still something that's occurring
  3016. 1:33:26linearly. And because of linear
  3017. 1:33:28constraints, you know, unless you are
  3018. 1:33:30using uh I don't know, like a a Gemini
  3019. 1:33:32flash model or you're using fast mode
  3020. 1:33:34and burning through your claw token uh
  3021. 1:33:36usage limits, this is going to take a
  3022. 1:33:37fair amount of time.
  3023. 1:33:38This process over here, literally to
  3024. 1:33:40just like launch the browser, could take
  3025. 1:33:425 seconds. This process to navigate to
  3026. 1:33:43the website could take 5 seconds. Taking
  3027. 1:33:45a page screenshot could take 15 seconds.
  3028. 1:33:47Identifying the contact form could take
  3029. 1:33:49a minute. You know, if you stack it all
  3030. 1:33:50up, basically what's occurring is this
  3031. 1:33:52whole process here
  3032. 1:33:54might take literally 2 to 3 minutes
  3033. 1:33:57per form if you're operating naively
  3034. 1:33:59using a slower model. And if you're
  3035. 1:34:00operating non-naively, if you're using a
  3036. 1:34:02smarter model, then obviously you have
  3037. 1:34:04to weigh that against cost and and and
  3038. 1:34:05token usage and stuff like that. So, I
  3039. 1:34:07don't know. Let's hypothetically say, in
  3040. 1:34:09my case, I wanted to reach out to
  3041. 1:34:11you know, 1,000 people.
  3042. 1:34:14Well, if it takes me two to three
  3043. 1:34:15minutes a form, that's 1,000 * 2. That's
  3044. 1:34:182,000 minutes, which divided by 60 is
  3045. 1:34:20like 30 hours or something like that,
  3046. 1:34:22right? That's a very long time. It's
  3047. 1:34:24going to take me a whole day.
  3048. 1:34:25So, instead of just doing one agent,
  3049. 1:34:27what I'm going to do is I'm basically
  3050. 1:34:29going to give every agent its own both
  3051. 1:34:31Chrome instance and then even its own a
  3052. 1:34:32workspace and then open up its autonomy
  3053. 1:34:35so that it can make some advanced
  3054. 1:34:36decisions to basically help it build its
  3055. 1:34:38own tooling if it needs in order to like
  3056. 1:34:40navigate website pages or whatever. Now,
  3057. 1:34:41what this is going to look like is
  3058. 1:34:42pretty similar to our previous
  3059. 1:34:45you know, stochastic multi-agent
  3060. 1:34:47consensus prompt. We're basically we
  3061. 1:34:49have a user up top, okay? And this is
  3062. 1:34:51us.
  3063. 1:34:52And what we're going to do is we're
  3064. 1:34:54going to give all the context about our
  3065. 1:34:57task, whatever it is that we want, you
  3066. 1:34:58know, fill out a form or I don't know,
  3067. 1:35:00do some lead gen to an orchestrator
  3068. 1:35:02agent, which in this case I'll do Claude
  3069. 1:35:05and we'll just do Opus.
  3070. 1:35:07Which
  3071. 1:35:08in my case is going to be 4.6, maybe in
  3072. 1:35:09your case it's a better model.
  3073. 1:35:11And then what that will do is it'll
  3074. 1:35:12spawn and set up, you know, however many
  3075. 1:35:15agents we want in separate windows. That
  3076. 1:35:18will then all in parallel navigate to
  3077. 1:35:20the site, find the form, fill the
  3078. 1:35:21fields, and then do the submission.
  3079. 1:35:23And so, basically, instead of it taking
  3080. 1:35:24two minutes per form, we can do is we
  3081. 1:35:26can actually submit, you know, however
  3082. 1:35:27many forms. So, I don't know. Let's say
  3083. 1:35:28we have like 10 agents. We'd submit 10
  3084. 1:35:31forms in the same amount of time it took
  3085. 1:35:32to submit uh one. So, maybe for us it'll
  3086. 1:35:34be 120 seconds.
  3087. 1:35:36And then what we do is we just we just
  3088. 1:35:37increase this as necessary. I mean, I
  3089. 1:35:39could theoretically have 500 operating
  3090. 1:35:40if I had the computing power. So, you
  3091. 1:35:42know, if previously it was one form
  3092. 1:35:45in
  3093. 1:35:46uh what did I say? Two minutes.
  3094. 1:35:48That means the form per minute rate is
  3095. 1:35:50like 0.5 forms a minute, right?
  3096. 1:35:53But now if we spin up 10, and we do 10
  3097. 1:35:56in 2 minutes, we're up to five a minute.
  3098. 1:35:59If we spin up, I don't know, 100,
  3099. 1:36:01then we're up to 50 a minute.
  3100. 1:36:03And you know, if my goal was 2,000 a day
  3101. 1:36:05and we're at 50 a minute, then obviously
  3102. 1:36:062,000 divided by 50 means we can get
  3103. 1:36:08this whole thing done in 40 minutes. And
  3104. 1:36:10you know, depending on the the the list
  3105. 1:36:12and whatever the heck you got, obviously
  3106. 1:36:13the constraints change. But um this is
  3107. 1:36:15how you can have multiple Chrome
  3108. 1:36:17instances operating simultaneously, it
  3109. 1:36:19navigating the website and stuff like
  3110. 1:36:21that. What I have here is a skill called
  3111. 1:36:23multi-agent-chrome.
  3112. 1:36:25And again, this is something you can
  3113. 1:36:26implement using whatever um context
  3114. 1:36:29framework you want, whether it's a
  3115. 1:36:30skill, whether it's like a Claude Gemini
  3116. 1:36:32or Agent 7D, whatever the heck you want.
  3117. 1:36:34What this basically forces it to do is
  3118. 1:36:35to orchestrate parallel browser
  3119. 1:36:36automation using multiple Chrome
  3120. 1:36:38DevTools MCP instances. And this is used
  3121. 1:36:41when a task requires doing the same
  3122. 1:36:42browser action across many targets
  3123. 1:36:44simultaneously. So some good examples
  3124. 1:36:46are submitting forms, filling apps,
  3125. 1:36:47scraping pages that need JavaScript
  3126. 1:36:49rendering, and and whatever.
  3127. 1:36:50And so what's occurring down here is
  3128. 1:36:52basically we have a top-level business
  3129. 1:36:54workspace, which is sort of the folder
  3130. 1:36:56that I'm in right now. And this actually
  3131. 1:36:58interacts with a bunch of Chrome agents,
  3132. 1:37:00which all have their own little MCP
  3133. 1:37:03servers, their own little um quad.mds,
  3134. 1:37:06and so on and so forth.
  3135. 1:37:07And then they all communicate with a
  3136. 1:37:09centralized chat.
  3137. 1:37:10And if they run into problems on
  3138. 1:37:12websites, if they have any reports they
  3139. 1:37:13want to give, basically what happens is
  3140. 1:37:15this orchestrator just checks the chat
  3141. 1:37:17every 30 seconds or so, okay?
  3142. 1:37:19So the very first step is it determines
  3143. 1:37:20how many agents are needed, then it
  3144. 1:37:22launches all the Chrome instances, and
  3145. 1:37:23resets the chat file because uh you
  3146. 1:37:25know, previous runs may have that. And
  3147. 1:37:27then you can see how every individual
  3148. 1:37:29sub-agent actually monitors its own um
  3149. 1:37:31task list by basically just pumping
  3150. 1:37:33things into a chat.
  3151. 1:37:35This is one of the simplest and easiest
  3152. 1:37:36ways of getting this specific design
  3153. 1:37:38pattern done. As mentioned, you guys can
  3154. 1:37:39get this down below if you want, but I'm
  3155. 1:37:41just going to give you guys a simple
  3156. 1:37:42example, which in my case is going to be
  3157. 1:37:43just finding Vancouver rentals because
  3158. 1:37:45I'm, you know, considering getting a
  3159. 1:37:46rental um down there. And so, you know,
  3160. 1:37:49rather than have it give you crappy
  3161. 1:37:51results, this thing can actually
  3162. 1:37:52navigate like Craigslist, Facebook
  3163. 1:37:54Marketplace, Kijiji, whatever the heck
  3164. 1:37:55you want. Uh and the the specific script
  3165. 1:37:57is right over here. So, hypothetically,
  3166. 1:37:59what it'll do is it'll just open up a
  3167. 1:38:00new window. And then I'll go over here
  3168. 1:38:02and then I'll write uh I want to find
  3169. 1:38:05a rental in Vancouver, needs to be
  3170. 1:38:0815-minute walk from the
  3171. 1:38:11Granville SkyTrain station downtown. Use
  3172. 1:38:15multi-agent
  3173. 1:38:16Chrome to navigate
  3174. 1:38:18through sites and give me
  3175. 1:38:21high-quality sleek places
  3176. 1:38:24under 2.5, let's say
  3177. 1:38:272K to 2.5K.
  3178. 1:38:29Other restrictions, like one bed, one
  3179. 1:38:33bath.
  3180. 1:38:34Reasonably
  3181. 1:38:35near the water. Needs AC built in. Okay.
  3182. 1:38:39So, I'm giving it a high-level um, you
  3183. 1:38:42know, piece of instruction. And sorry,
  3184. 1:38:43what I meant to do is actually do prompt
  3185. 1:38:45contract after this. And now I want it
  3186. 1:38:47to give me like a very clear contract.
  3187. 1:38:49So, it's going to give me a list of five
  3188. 1:38:50to 10 rental apartments. Why don't we
  3189. 1:38:52say 20 rental apartments?
  3190. 1:38:55And then I'll say 1.2 km is fine.
  3191. 1:38:58We'll say near water south of Drake or
  3192. 1:39:00west of Burrard.
  3193. 1:39:02Okay, I'm just going to make some
  3194. 1:39:03changes here. And then I'll say that
  3195. 1:39:05sounds pretty good. Go for it. And now
  3196. 1:39:06it's going to actually launch the
  3197. 1:39:08multi-agent Chrome scraping. So, it's
  3198. 1:39:10then going to invoke the skill. I'm just
  3199. 1:39:12going to keep my hands off.
  3200. 1:39:14What it'll do next is actually spawn um
  3201. 1:39:16four [clears throat] parallel Chrome
  3202. 1:39:17agents, one per rental site. So, it'll
  3203. 1:39:19determine that there are four rental
  3204. 1:39:20sites that it's going to be running
  3205. 1:39:22through. And uh it'll just have one
  3206. 1:39:24Chrome instance sort of do everything
  3207. 1:39:25there.
  3208. 1:39:26Uh per site. So, now we have the four
  3209. 1:39:28instances. I'm just going to open this
  3210. 1:39:29up here. Open this up here. I'll move
  3211. 1:39:33this one down over here. And I'll also
  3212. 1:39:35move this one down over here.
  3213. 1:39:37Obviously, you could use an approach
  3214. 1:39:38like this for pretty nefarious purposes.
  3215. 1:39:40Um so, you do have to be cognizant of
  3216. 1:39:42that that a lot of people and websites
  3217. 1:39:44are probably, um, you know, they're
  3218. 1:39:46looking to verify whether or not you are
  3219. 1:39:47a person. And so there are multiple
  3220. 1:39:50things you can do to get around that if
  3221. 1:39:51you so wanted to, like using custom, um,
  3222. 1:39:54browser fingerprinting and whatnot. And
  3223. 1:39:56I think that's a story for another
  3224. 1:39:57course because I don't really want this
  3225. 1:39:59course to be accused of showing you guys
  3226. 1:40:00how to spin up like 500 Chrome instances
  3227. 1:40:02scraping all sorts of illicit
  3228. 1:40:03information on the internet using unique
  3229. 1:40:05browser fingerprints. But, uh, that
  3230. 1:40:06stuff is definitely possible and there
  3231. 1:40:08are probably like a lot of people doing
  3232. 1:40:10stuff similar to this right now that are
  3233. 1:40:11just way farther ahead in terms of
  3234. 1:40:12their,
  3235. 1:40:13uh, you know, understanding of agents
  3236. 1:40:15and stuff like that. Now, after the
  3237. 1:40:1630-second or so wait time, these will
  3238. 1:40:18receive their instructions and they'll
  3239. 1:40:19actually check the main thread and then
  3240. 1:40:20they'll load in their websites. So, this
  3241. 1:40:22one up here spawned, uh, lib. And I've
  3242. 1:40:24never used that site before. This one's
  3243. 1:40:26padmapper.com, which is another one. You
  3244. 1:40:28know, these are all like websites and
  3245. 1:40:30resources I probably would not have
  3246. 1:40:31looked at. And as a result, I'm going to
  3247. 1:40:33get some more of like a search spread.
  3248. 1:40:35I'm going to, again, for a big chunk of
  3249. 1:40:37the search space much faster than if I
  3250. 1:40:39were to have done all this stuff
  3251. 1:40:40manually. What's cool is these consume
  3252. 1:40:42directly into pages for me. These can
  3253. 1:40:44click on links and stuff like that.
  3254. 1:40:45They're obviously modifying filters and
  3255. 1:40:47and whatnot autonomously so that they're
  3256. 1:40:49not just getting a bunch of bogus
  3257. 1:40:50results. And at the end I get a
  3258. 1:40:51high-quality filtered list of apartments
  3259. 1:40:53that are,
  3260. 1:40:54you know, within my specifications.
  3261. 1:40:56Okay, I just turned my camera up because
  3262. 1:40:57I wanted some additional room in the
  3263. 1:40:58bottom left-hand side to really drill a
  3264. 1:41:00few important points home. The first is
  3265. 1:41:02your context window. Now, remember
  3266. 1:41:05earlier how we talked about the
  3267. 1:41:06claude.md,
  3268. 1:41:08the gemini.md,
  3269. 1:41:10and then the agents.md.
  3270. 1:41:12Over here we just have Claude, but I
  3271. 1:41:14just want you to treat this as all, uh,
  3272. 1:41:16three of them.
  3273. 1:41:17That's not the only thing that gets
  3274. 1:41:19injected, so to speak, in your context.
  3275. 1:41:24You have a variety of other things.
  3276. 1:41:26Now, you have your system prompt up top.
  3277. 1:41:30You then have the claude.md, agents.md,
  3278. 1:41:32and whatever else.
  3279. 1:41:34You have a file, at least in Claude
  3280. 1:41:36code, called memory.md,
  3281. 1:41:39although there are analogs in other
  3282. 1:41:41coding platforms.
  3283. 1:41:44And then you also have skills and tools.
  3284. 1:41:49We've chatted a lot about MCP over the
  3285. 1:41:52course of the last hour and a half or
  3286. 1:41:54so, right? Well, MCP is a type of skill
  3287. 1:41:57and tool.
  3288. 1:41:58We also have the actual skills
  3289. 1:42:00themselves. So, remember the agent
  3290. 1:42:02reviewer when we were doing sub-agent
  3291. 1:42:04verification loops?
  3292. 1:42:06Well, that was an example of a skill.
  3293. 1:42:09You think to the prompt contracts,
  3294. 1:42:12those were examples of skills.
  3295. 1:42:16And the reason why I'm going into depth
  3296. 1:42:18here is because each of these sections
  3297. 1:42:21can consume a tremendous number of
  3298. 1:42:23tokens.
  3299. 1:42:25And you're not given an unlimited number
  3300. 1:42:26of tokens to start with. Everything in
  3301. 1:42:28life is finite, including uh you know,
  3302. 1:42:31your Claude or your Gemini context
  3303. 1:42:33window.
  3304. 1:42:34Now, most models right now are somewhere
  3305. 1:42:36between
  3306. 1:42:38uh I don't know if it's like 4.6, we're
  3307. 1:42:40probably talking 200k to 1 million.
  3308. 1:42:44If we're talking Gemini, you know, we
  3309. 1:42:46have like 3.1 and there there are a
  3310. 1:42:48couple other ones, obviously. But by the
  3311. 1:42:50time you guys are watching this, they'll
  3312. 1:42:51probably be more.
  3313. 1:42:53And then, you know, you have uh GPT 5.4
  3314. 1:42:57and then Codex 5.3, but the 5.4 is
  3315. 1:42:59coming out.
  3316. 1:43:00You know, most models nowadays have
  3317. 1:43:03somewhere in the realm of between 200k
  3318. 1:43:05to 1 million.
  3319. 1:43:07And to be clear, um a token is not a
  3320. 1:43:09word.
  3321. 1:43:11A token is about
  3322. 1:43:140.7
  3323. 1:43:16words.
  3324. 1:43:17So, if you think about it in that vein,
  3325. 1:43:18what this means is these 200k tokens
  3326. 1:43:21sort of actually equate to somewhere
  3327. 1:43:23between like 140,000 words to about
  3328. 1:43:26700,000 words, okay?
  3329. 1:43:28Um but this context window obviously it
  3330. 1:43:30filled up the more that you talk with
  3331. 1:43:32it. And unfortunately, one common and
  3332. 1:43:35major problem in large language models,
  3333. 1:43:38specifically the types that we're
  3334. 1:43:39dealing with in this course, are as time
  3335. 1:43:42goes on and you talk to it more and more
  3336. 1:43:45and more,
  3337. 1:43:47what you find is the average quality
  3338. 1:43:51goes down.
  3339. 1:43:53So, quality
  3340. 1:43:55as a factor or a byproduct of token
  3341. 1:43:59count, typically starts pretty high up
  3342. 1:44:01here at maybe, I don't know, 100%.
  3343. 1:44:05And then the longer and longer and
  3344. 1:44:07longer the number of tokens in your
  3345. 1:44:08context, the lower and lower and lower
  3346. 1:44:11the quality gets.
  3347. 1:44:12So, maybe this is a 10K.
  3348. 1:44:15Maybe this is at 50K.
  3349. 1:44:18Maybe this over here is a 200K.
  3350. 1:44:20And what that means is let's just
  3351. 1:44:22hypothetically say you're at your
  3352. 1:44:23199,000th
  3353. 1:44:26token, okay?
  3354. 1:44:28That means
  3355. 1:44:29that on a equivalent query that you
  3356. 1:44:33might have previously scored 100% at,
  3357. 1:44:37at, I don't know, 5 or 10K tokens,
  3358. 1:44:39at 199,000 tokens, you might only score
  3359. 1:44:4240% at.
  3360. 1:44:43Now, these numbers I basically pulled
  3361. 1:44:45out of my ass to be clear, but the point
  3362. 1:44:48I'm trying to make is the longer the
  3363. 1:44:49token count, basically the bigger the
  3364. 1:44:52context
  3365. 1:44:53length,
  3366. 1:44:55the lower the performance of the model.
  3367. 1:44:57And so, understanding context windows
  3368. 1:44:59and then learning a little bit of
  3369. 1:45:01context management, ways to proactively
  3370. 1:45:02manage all of these things, some of
  3371. 1:45:04which you have control over and some of
  3372. 1:45:06other things which you don't, is very
  3373. 1:45:08important.
  3374. 1:45:09It's also important, of course, because
  3375. 1:45:10of billing.
  3376. 1:45:12The more tokens that you use up,
  3377. 1:45:14obviously the more money that you spend.
  3378. 1:45:17And so, not only is it best from a
  3379. 1:45:18quality perspective over here to try and
  3380. 1:45:22push to the left side of this graph as
  3381. 1:45:23much as humanly possible, it's It's very
  3382. 1:45:25relevant from a financial perspective
  3383. 1:45:28over here because obviously the more
  3384. 1:45:30tokens that you use the more money you
  3385. 1:45:31spend. Case in point, just to make this
  3386. 1:45:33video I've spent something around $500
  3387. 1:45:36or so in tokens. Now that's because I'm
  3388. 1:45:38using a particular agent's fast mode
  3389. 1:45:41which bills me directly instead of just
  3390. 1:45:43using a monthly plan. But the point
  3391. 1:45:45remains, any sort of serious AI agent
  3392. 1:45:47application will start spending and
  3393. 1:45:49using a fair amount of your money.
  3394. 1:45:51Okay, and just before I move on, I want
  3395. 1:45:53to talk a tiny bit about the differences
  3396. 1:45:56between each of these in length. If I
  3397. 1:45:58open up an actual Claude instance here,
  3398. 1:46:01I open up one of these and then I go to
  3399. 1:46:03terminal
  3400. 1:46:04which is the current best way to
  3401. 1:46:06visualize this. And if I just maximize
  3402. 1:46:09this panel size, let's just make this as
  3403. 1:46:11big as humanly possible,
  3404. 1:46:13then I go {slash} context,
  3405. 1:46:15Claude will show us all of the things
  3406. 1:46:18currently consuming its tokens.
  3407. 1:46:20I'm going to zoom in here to make it
  3408. 1:46:21really really clear what's going on.
  3409. 1:46:25This right over here is your context
  3410. 1:46:27usage. And as you can see, they've
  3411. 1:46:29illustrated this as sort of a series of
  3412. 1:46:32squares where every square is, I don't
  3413. 1:46:34know, let's see, 1 2 3 4 5 6 7 8 9 10.
  3414. 1:46:37Okay, every square is, I think, 2,000
  3415. 1:46:40tokens or so.
  3416. 1:46:41And so what we're seeing is, despite the
  3417. 1:46:44fact that we have put no conversation
  3418. 1:46:48tokens, so we haven't spent any tokens
  3419. 1:46:51at all on conversation,
  3420. 1:46:53we're still
  3421. 1:46:55at 9,000
  3422. 1:46:57used.
  3423. 1:46:59You're probably wondering where the hell
  3424. 1:47:00are these 9,000 coming from? Are they
  3425. 1:47:01shadow billing me to try and
  3426. 1:47:04rinse my wallet as much as humanly
  3427. 1:47:06possible?
  3428. 1:47:07Well, a little bit. I mean, this system
  3429. 1:47:09prompt here, okay, which is partially
  3430. 1:47:12composed by your agents, your Gemini or
  3431. 1:47:14your Claude and MD and partially a few
  3432. 1:47:15additional things, it's actually already
  3433. 1:47:17consuming 4,900 tokens. So 2.5% of my
  3434. 1:47:21entire token count before I even send a
  3435. 1:47:22message is being used by in this case
  3436. 1:47:25probably the Claude.md.
  3437. 1:47:29But in addition you have other things
  3438. 1:47:30like memory files which are consuming
  3439. 1:47:312,000 tokens, okay? And the way that
  3440. 1:47:35they do that in Claude code is they use
  3441. 1:47:37something called a memory.md which
  3442. 1:47:39stores your preferences and some some
  3443. 1:47:41previous high-level things.
  3444. 1:47:43Then next up we have skills which are
  3445. 1:47:44consuming 1,700 tokens. What are these
  3446. 1:47:47skills? Well, you guys remember when we
  3447. 1:47:49made a bunch over here? If I go to the
  3448. 1:47:51top left-hand corner where it says doc
  3449. 1:47:52Claude skills, you know, every agenda
  3450. 1:47:55coding platform has their own
  3451. 1:47:56configuration for this stuff. But the
  3452. 1:47:58way that it works in Claude code is, you
  3453. 1:48:00know, you organize these workflows into
  3454. 1:48:01these skills.
  3455. 1:48:03Well, guess what? These skills aren't
  3456. 1:48:04free. In order for Claude to be able to
  3457. 1:48:07use these skills, okay? This multi-agent
  3458. 1:48:10orchestrator, it needs to store all
  3459. 1:48:12these tokens somewhere and then give it
  3460. 1:48:14to the model and that's what's going on
  3461. 1:48:15over here, right? Now the actual
  3462. 1:48:17messages that we've used
  3463. 1:48:19are only at eight tokens. And guess
  3464. 1:48:20what? That's actually, I think it's um
  3465. 1:48:22this word here context usage.
  3466. 1:48:26I'm not entirely sure, but I think
  3467. 1:48:28context usage just because of the way
  3468. 1:48:29that it's broken down or maybe context
  3469. 1:48:30usage plus this term here {slash}
  3470. 1:48:32context, you know, is equal to eight
  3471. 1:48:34tokens.
  3472. 1:48:35But know that, you know, that's that's
  3473. 1:48:37it. That's it for our whole token count.
  3474. 1:48:39So the other 158,000 of our 200,000
  3475. 1:48:41limit is currently free.
  3476. 1:48:43And so I mean this is a quick and easy
  3477. 1:48:45way obviously to visualize it inside of
  3478. 1:48:46Claude code, but um other platforms have
  3479. 1:48:48their own visualization mechanisms. Now
  3480. 1:48:50next up are these MCP tools. A way to
  3481. 1:48:52look at these MCP {underscore}
  3482. 1:48:54{underscore} Chrome dev tools
  3483. 1:48:56{underscore} {underscore} click. What is
  3484. 1:48:58that? Well, this is the tool that allows
  3485. 1:49:01Chrome to click on parts of the page.
  3486. 1:49:05Remember earlier when we were building
  3487. 1:49:07that little N8N flow?
  3488. 1:49:09Well, we were doing it by clicking on
  3489. 1:49:10various parts of the page.
  3490. 1:49:12How about drag, right? We can drag
  3491. 1:49:15things, get console message. These are
  3492. 1:49:17all basically buttons in some colossal
  3493. 1:49:21spaceship. Basically, we are in the
  3494. 1:49:24cockpit with Claude code and we're
  3495. 1:49:26telling it to do stuff for us. We don't
  3496. 1:49:28know what these buttons are. It does
  3497. 1:49:31because it's, you know, the ship
  3498. 1:49:32technician or the navigator or whatever.
  3499. 1:49:34And so it's clicking these buttons left,
  3500. 1:49:36right, and center for us to do various
  3501. 1:49:37things. That's how you can conceptualize
  3502. 1:49:39all these MCP tools and all these
  3503. 1:49:40skills. And what I really like about
  3504. 1:49:41this is it breaks everything down. So
  3505. 1:49:42here are memory files, okay, which I
  3506. 1:49:44talked about the memory.md, the
  3507. 1:49:46claude.md. You can see this is being
  3508. 1:49:48contributed to in a variety of ways. We
  3509. 1:49:50have a global claude.md, which is sort
  3510. 1:49:52of like a very high-level one with some
  3511. 1:49:53sparse instructions. We have a local
  3512. 1:49:55claude.md. We actually have the memory
  3513. 1:49:57down here and then we have all the
  3514. 1:49:58skills. You know, what was really
  3515. 1:49:59telling about that is despite the fact
  3516. 1:50:01that in this diagram conversation
  3517. 1:50:02history is like the biggest chunk of it
  3518. 1:50:04all. Notice that in reality,
  3519. 1:50:06conversation history for us, at least at
  3520. 1:50:08the beginning, was nothing. You know, it
  3521. 1:50:10was actually a tremendous number of
  3522. 1:50:12tokens, about 10% of our entire context
  3523. 1:50:14window, dedicated to just this little
  3524. 1:50:16chunk.
  3525. 1:50:17And that is a problem because if you're
  3526. 1:50:20not careful, your claude.md with all
  3527. 1:50:22that's rules are going to get really,
  3528. 1:50:23really long. Your memory.md with all
  3529. 1:50:26your preferences is going to get huge.
  3530. 1:50:27Same thing with all the skills and tools
  3531. 1:50:29and stuff like that. And your
  3532. 1:50:29conversation history, when you actually
  3533. 1:50:31do get to conversating with the model,
  3534. 1:50:33will be very, very, very small.
  3535. 1:50:35You know, instead of starting, I don't
  3536. 1:50:37know, somewhere in this region here,
  3537. 1:50:40you're actually, because the context
  3538. 1:50:42window of all of your BS is so big, you
  3539. 1:50:44might actually like start in the
  3540. 1:50:45effective area over here.
  3541. 1:50:47And obviously this is not what you want
  3542. 1:50:49to begin an agent conversation at
  3543. 1:50:51because if you start here, then it's
  3544. 1:50:52obviously only downhill from there. Now
  3545. 1:50:54this takes me to the logical question
  3546. 1:50:56of, you know, hey Nick, what happens
  3547. 1:50:58when you run out of context? Because
  3548. 1:51:00obviously that's going to happen.
  3549. 1:51:02Well, when there are 50,000 tokens,
  3550. 1:51:05let's say, out of 200,000, okay? No
  3551. 1:51:07problem. You are having the full
  3552. 1:51:09conversation history with the model and
  3553. 1:51:12the model basically gets literally every
  3554. 1:51:14message starting from message number
  3555. 1:51:15one, two, three, four, all the way down
  3556. 1:51:18to I don't know, message number 25. And
  3557. 1:51:20then what it does is it takes all of
  3558. 1:51:22this context and then feeds it into its
  3559. 1:51:24big neural network to generate message
  3560. 1:51:26number 26, right?
  3561. 1:51:28However, when we get to a certain
  3562. 1:51:29length, right? I don't know, let's say
  3563. 1:51:31message 50. Obviously, there's no
  3564. 1:51:33there's no tokens left anymore in the
  3565. 1:51:34context window. Maybe we're like 199
  3566. 1:51:36Well, not actually, but probably be like
  3567. 1:51:37155k
  3568. 1:51:39out of 200k.
  3569. 1:51:40Well, what happens is all models now
  3570. 1:51:42have some sort of what's called like
  3571. 1:51:44auto compact limit.
  3572. 1:51:46I mean, basically all of them have
  3573. 1:51:47adopted this convention, where when the
  3574. 1:51:50number of tokens that you're using,
  3575. 1:51:52let's say the
  3576. 1:51:54limit is right over here.
  3577. 1:51:57When the number of tokens that you're
  3578. 1:51:58using gets to this point, okay?
  3579. 1:52:01You know, it fills up,
  3580. 1:52:03then fills up,
  3581. 1:52:04fills up,
  3582. 1:52:06fills up, and fills up.
  3583. 1:52:08What happens is this triggers a
  3584. 1:52:10mechanism called compaction,
  3585. 1:52:12where we take all of the information
  3586. 1:52:15here,
  3587. 1:52:18which is I don't know,
  3588. 1:52:19maybe like 80% or so of the whole
  3589. 1:52:21context, and then we compress it. I want
  3590. 1:52:24you to imagine right now there's like a
  3591. 1:52:25a big hydraulic press type thing over
  3592. 1:52:28here,
  3593. 1:52:29and it's pushing all of this context
  3594. 1:52:31down, it's squishing it. Basically,
  3595. 1:52:33what's going to occur
  3596. 1:52:34is we're going to erase the vast
  3597. 1:52:36majority of this,
  3598. 1:52:38and then, you know, instead of consuming
  3599. 1:52:4080%, we're going to cram all that
  3600. 1:52:42information into maybe, I don't know, 30
  3601. 1:52:44or 40% or so.
  3602. 1:52:47And so that is called compaction,
  3603. 1:52:50and it occurs on a relatively regular
  3604. 1:52:52basis across the model ecosystem.
  3605. 1:52:55The issue with the compaction, or you
  3606. 1:52:57know, compression, or whatever the heck
  3607. 1:52:58you want to call it, it's the same idea,
  3608. 1:53:00is during this summarization and like
  3609. 1:53:03densification process, we are going to
  3610. 1:53:05drop outputs from tools. We are going to
  3611. 1:53:08remove some information and context that
  3612. 1:53:11might actually be useful to you now at,
  3613. 1:53:13you know, message 50 from message four,
  3614. 1:53:15which might actually like eliminate a
  3615. 1:53:16mistake. And so, because of that, you
  3616. 1:53:18know, you are going to lose some of the
  3617. 1:53:19quality.
  3618. 1:53:20Um the benefit is obviously you will
  3619. 1:53:22significantly improve the information
  3620. 1:53:24density. And what do I mean by
  3621. 1:53:25information density? I mean literally
  3622. 1:53:26like if it's like um hello,
  3623. 1:53:30how are
  3624. 1:53:32you doing?
  3625. 1:53:35I mean, let's say somewhere in your
  3626. 1:53:36context you have the term you the
  3627. 1:53:38sentence hello, how are you doing? Well,
  3628. 1:53:40hello is actually two tokens. How is
  3629. 1:53:42one, are is one, you is one, doing is
  3630. 1:53:46three along the question mark. So, in
  3631. 1:53:48total, if you count all the stuff up,
  3632. 1:53:50depending on the uh thing you're using
  3633. 1:53:51to do the embedding, that's uh eight,
  3634. 1:53:53right?
  3635. 1:53:54Well, you know, compaction is literally
  3636. 1:53:57going to take this sentence and then
  3637. 1:53:58it's going to compress it. So, it's
  3638. 1:53:59going to say hi,
  3639. 1:54:01how are you?
  3640. 1:54:04So, now instead of that's one token
  3641. 1:54:06here, one token here, one token here,
  3642. 1:54:08one token here equals four.
  3643. 1:54:10It will literally get all of the context
  3644. 1:54:12that it can and then try and squish it
  3645. 1:54:14so that the same meaning is available in
  3646. 1:54:16fewer tokens and fewer words wherever
  3647. 1:54:18possible. And then it'll just run this
  3648. 1:54:20naively across your entire um context.
  3649. 1:54:23And so, this is going to occur every
  3650. 1:54:24single time. Obviously, it's something
  3651. 1:54:25that we want to avoid occurring if we
  3652. 1:54:27have sensitive and important data, but
  3653. 1:54:29you know, it does allow us to continue
  3654. 1:54:31conversating with models. Before we had
  3655. 1:54:33context compression in some form of auto
  3656. 1:54:34comp action, um basically we would just
  3657. 1:54:37run out of tokens and we'd have to
  3658. 1:54:38restart a totally new session. So, this
  3659. 1:54:40just does sort of that intermediate step
  3660. 1:54:41that most people were doing before where
  3661. 1:54:42they would like like take all of that,
  3662. 1:54:44you know, try and summarize it in some
  3663. 1:54:47other model and then paste it back into
  3664. 1:54:48uh into into another one. Now, that
  3665. 1:54:50takes us to what most model
  3666. 1:54:52practitioners are using nowadays, which
  3667. 1:54:54is a variant of something that uh people
  3668. 1:54:57call the iceberg technique.
  3669. 1:54:59Now, in case you guys have never seen an
  3670. 1:55:01iceberg before, I'm Canadian, so we have
  3671. 1:55:03them everywhere, including literally in
  3672. 1:55:05the river across the street from my
  3673. 1:55:07house.
  3674. 1:55:08Not uh actual icebergs, but little ice
  3675. 1:55:10floats.
  3676. 1:55:11The way that they work is basically you
  3677. 1:55:13have um a section that's visible above,
  3678. 1:55:16which is usually quite massive and quite
  3679. 1:55:17intimidating. And then you're like, "Oh
  3680. 1:55:19my god, that's a really big iceberg."
  3681. 1:55:21And then what you don't realize is
  3682. 1:55:22underneath the iceberg is actually like
  3683. 1:55:24two or three times as big.
  3684. 1:55:26And so above, the stuff that's like
  3685. 1:55:28immediately visible,
  3686. 1:55:30or in our terms for the model,
  3687. 1:55:33accessible immediately,
  3688. 1:55:36is actually only a very small percentage
  3689. 1:55:38of the total iceberg.
  3690. 1:55:42And so in the context of our model, what
  3691. 1:55:44we store above, which is immediately
  3692. 1:55:46accessible, it's the sort of the stuff
  3693. 1:55:47that's like visible to the plain eye, is
  3694. 1:55:50we'll store our memory,
  3695. 1:55:52we'll store our Claude or our agents or
  3696. 1:55:54our Gemini.md,
  3697. 1:55:56we'll store our local memory as well.
  3698. 1:55:59There's different types, there's global,
  3699. 1:56:00local. We'll store all of the current
  3700. 1:56:02task contexts, everything that our tools
  3701. 1:56:05are doing, then maybe any active file
  3702. 1:56:07contents, okay? And so all of this stuff
  3703. 1:56:09here is like basically always accessible
  3704. 1:56:11to us. It's literally just in our
  3705. 1:56:12prompt.
  3706. 1:56:13But then, what people are doing to
  3707. 1:56:15reduce the total number of tokens that
  3708. 1:56:16they require, is they're abstracting
  3709. 1:56:18away everything else. And then they're
  3710. 1:56:19just making it accessible to the model
  3711. 1:56:21if it needs it.
  3712. 1:56:22What I mean by that is instead of
  3713. 1:56:24putting all of the files in your code
  3714. 1:56:26base in a prompt, what it does is it
  3715. 1:56:28gives you a tool called read. What read
  3716. 1:56:30can do is read can at any time read a
  3717. 1:56:32file, okay? But instead of putting the
  3718. 1:56:35entire file in there, for now all it
  3719. 1:56:36does is just puts the titles.
  3720. 1:56:38So if you say, "Hey, I want you to grab
  3721. 1:56:40the um information on the iceberg
  3722. 1:56:42technique." And then in your workspace
  3723. 1:56:44you have a file called iceberg
  3724. 1:56:45technique.md,
  3725. 1:56:46it'll know it doesn't actually have to
  3726. 1:56:47read all of the files in your workspace,
  3727. 1:56:49it only has to read iceberg.md, right?
  3728. 1:56:52Same thing with the full code base. You
  3729. 1:56:53have tools like grep and glob. And these
  3730. 1:56:56tools are sort of analogous. Instead of
  3731. 1:56:58reading a whole file, what this does is
  3732. 1:57:00allows you to hone in on a specific
  3733. 1:57:01segment of text. You know, if this is my
  3734. 1:57:03entire code file, okay, and
  3735. 1:57:05hypothetically, let's just say it's
  3736. 1:57:07really big and, you know, there's
  3737. 1:57:09there's a lot of stuff. But the only
  3738. 1:57:11thing that I actually care about is this
  3739. 1:57:13little segment over here,
  3740. 1:57:16then why would I load all of the, I
  3741. 1:57:19don't know, 10K tokens?
  3742. 1:57:21I don't need to, okay? Realistically,
  3743. 1:57:24what I can do as a smart model is I
  3744. 1:57:26could use grep and glob, these tools, to
  3745. 1:57:29maybe hone in only on this segment over
  3746. 1:57:31here, which is, I don't know, let's just
  3747. 1:57:33say 2K tokens.
  3748. 1:57:35And because it contains some of the text
  3749. 1:57:37before and some of the text after, you
  3750. 1:57:39know, this still usually gives us enough
  3751. 1:57:40context to
  3752. 1:57:42um tell the model what it needs in order
  3753. 1:57:44to to finish its function. Then you also
  3754. 1:57:45have web data via web fetch. Web data's
  3755. 1:57:48pretty cool. You can kind of think of it
  3756. 1:57:50as the same thing. Obviously, it doesn't
  3757. 1:57:51have access to the whole internet, but
  3758. 1:57:52it can make search queries, right? And
  3759. 1:57:54so, because it's able to use some
  3760. 1:57:56general reasoning, when you say, "Hey,
  3761. 1:57:57what's the iceberg technique?" first
  3762. 1:57:59it's going to start by looking for file
  3763. 1:58:01contents called iceberg technique. If it
  3764. 1:58:03can't find any, maybe it'll quickly type
  3765. 1:58:05F iceberg to look through the code base.
  3766. 1:58:07If it can't find that, you know, maybe
  3767. 1:58:09it'll look through some other things
  3768. 1:58:10like uh memory files in the skills
  3769. 1:58:12library. But if it can't find that,
  3770. 1:58:14it'll say, "Okay, cool. So, we don't
  3771. 1:58:15have this in our in our context window
  3772. 1:58:17right now, and we don't even have access
  3773. 1:58:19to it within our space, but it's
  3774. 1:58:20probably somewhere on the internet, so
  3775. 1:58:21I'm just going to Google iceberg
  3776. 1:58:22technique." And then what it'll do is it
  3777. 1:58:25won't even take the entire thing. It'll
  3778. 1:58:26just grab top-level links, and then
  3779. 1:58:28it'll look at the URL, you know, uh and
  3780. 1:58:31the URL is about, guess what, icebergs.
  3781. 1:58:33I'm going off the map here, but
  3782. 1:58:35hopefully you guys understand what I
  3783. 1:58:36mean. Then if there's three URLs, one of
  3784. 1:58:38them's about icebergs, then instead of
  3785. 1:58:39reading all of them, it's only going to
  3786. 1:58:40read this one. And so, it's sort of like
  3787. 1:58:42a successive narrowing of the lens until
  3788. 1:58:45eventually it gets to, you know, what
  3789. 1:58:47you want. And it doesn't load the the
  3790. 1:58:50of all of the context, but it just has
  3791. 1:58:53the opportunity select all of this. You
  3792. 1:58:55know, it starts here, it goes here, goes
  3793. 1:58:57here, goes here, goes here, goes here,
  3794. 1:58:58and then goes here, and then finally
  3795. 1:58:59it's at its goal.
  3796. 1:59:01You can do the same thing in a variety
  3797. 1:59:02of other ways. You can use bash, get
  3798. 1:59:03history, and so on and so forth, but
  3799. 1:59:06um essentially what you want to do is
  3800. 1:59:08instead of storing all of the
  3801. 1:59:09information like the full code base, the
  3802. 1:59:11file contents, the web data, all the get
  3803. 1:59:13history, all the skills, and everything
  3804. 1:59:14like that. What you do is you just store
  3805. 1:59:16the ability to access it on demand.
  3806. 1:59:19And then inside the context, okay, the
  3807. 1:59:22tiny chunk that you do, I mean in this
  3808. 1:59:23diagram it says 10% 90%, you know, I
  3809. 1:59:25think in reality it's probably closer to
  3810. 1:59:272080 or maybe 3070.
  3811. 1:59:30Um here's where you store stuff that's
  3812. 1:59:31just like
  3813. 1:59:32it it needs to be the same all the time.
  3814. 1:59:34It always needs to be present. There
  3815. 1:59:36always needs to be some sort of patterns
  3816. 1:59:38that are learned, some sort of active
  3817. 1:59:39file contexts or
  3818. 1:59:41contexts or current task context. You
  3819. 1:59:43can think of this as the difference
  3820. 1:59:44between naive versus strategic context
  3821. 1:59:48loading. Now way back in the day, and
  3822. 1:59:50when I say way back in the day, I mean
  3823. 1:59:52like 2023. Good God, I'm getting old.
  3824. 1:59:55You know, when you were working with an
  3825. 1:59:56agent, you would dump in the whole code
  3826. 1:59:57base. You would honestly just copy and
  3827. 2:00:00paste everything and just hope to God
  3828. 2:00:01that it that it knew what it was doing.
  3829. 2:00:03Obviously, because this was
  3830. 2:00:05extraordinarily infeasible, there were
  3831. 2:00:08so many tokens in the context.
  3832. 2:00:11Tons of it was lost, and you routinely
  3833. 2:00:13ran out of context limits.
  3834. 2:00:15Well, nowadays what we've done is we
  3835. 2:00:16basically built in a whole tool stack
  3836. 2:00:18where instead of all of the file, you
  3837. 2:00:20just read selectively and only the
  3838. 2:00:22relevant functions.
  3839. 2:00:23You know, you have a cloud.net which is
  3840. 2:00:25basically a compression a compression
  3841. 2:00:27function which is stores your
  3842. 2:00:28preferences.
  3843. 2:00:29Skills instead of all being read, you
  3844. 2:00:31know, we only read a specific segment of
  3845. 2:00:33them. This is technically called a YAML
  3846. 2:00:35front matter, which is just a tiny
  3847. 2:00:37little section at the beginning of the
  3848. 2:00:38skill. You can actually see this if we
  3849. 2:00:40go back to the skill.net and I make this
  3850. 2:00:42visible, right? Only this section up
  3851. 2:00:44here is actually and know, can actually
  3852. 2:00:46see this up here on the YAML front
  3853. 2:00:48matter is this segment. Only this
  3854. 2:00:51segment is actually loaded into context
  3855. 2:00:53until you ask for, you know, more
  3856. 2:00:55information about create proposal. And
  3857. 2:00:57that's just because this little space
  3858. 2:00:58invader sees you it has the ability to
  3859. 2:01:00call a create proposal skill, but it
  3860. 2:01:02doesn't need to know all the rest of
  3861. 2:01:04this because what's realistically going
  3862. 2:01:05to happen well, 90% of the time you
  3863. 2:01:06won't even ask. There's so many other
  3864. 2:01:07skills that it'll probably be using.
  3865. 2:01:09We're also now doing things like uh
  3866. 2:01:10summarizing tool results. So, instead of
  3867. 2:01:13storing the entire thing in your
  3868. 2:01:15context, you know, we just store like a
  3869. 2:01:17very short summary of basically the
  3870. 2:01:18input output. So, way back in the day I
  3871. 2:01:20used to think that all of an agent was
  3872. 2:01:23really just its intelligence. It was
  3873. 2:01:25just the core model, right? Which in
  3874. 2:01:27this case would be Opus 4.6.
  3875. 2:01:30But what I've come to quickly realize is
  3876. 2:01:32although models themselves are quite
  3877. 2:01:34intelligent, it's really the
  3878. 2:01:36architecture that we've wrapped around
  3879. 2:01:37it. I want you to pretend this little
  3880. 2:01:39space invader is kind of now in I don't
  3881. 2:01:42know a
  3882. 2:01:43house of some kind. The house has a
  3883. 2:01:45little chimney with a fireplace. It has
  3884. 2:01:47a little place where it could I don't
  3885. 2:01:48know, roast a nice turkey. You know, it
  3886. 2:01:51has a nice bed that it can go to sleep
  3887. 2:01:52in every night. You know, this agent by
  3888. 2:01:55itself probably wouldn't last super long
  3889. 2:01:57out there on the savanna, but because
  3890. 2:01:58we've built all this infrastructure
  3891. 2:02:00around it, because we built roads for
  3892. 2:02:01it, we built ways for it to communicate
  3893. 2:02:03and stuff like that, you know, it's it's
  3894. 2:02:05capable of actually doing a lot of very
  3895. 2:02:06economically valuable work for us.
  3896. 2:02:09And so, just like human beings
  3897. 2:02:11way back in the day had to conceptualize
  3898. 2:02:13the idea of
  3899. 2:02:14I don't know, like a like a spear or
  3900. 2:02:16something to hunt um um
  3901. 2:02:19saber-tooth tigers on the plains, you
  3902. 2:02:22know, so too do these agents use tools
  3903. 2:02:26effectively to solve problems in their
  3904. 2:02:28environment and ultimately get us the
  3905. 2:02:30users what we want. Now, obviously
  3906. 2:02:32there's context window management like
  3907. 2:02:35we just talked about. And that's for
  3908. 2:02:37optimizing the usage of a specific
  3909. 2:02:40model, okay, in the choice of a model.
  3910. 2:02:43but there's also the ability to choose
  3911. 2:02:44different models for different purposes.
  3912. 2:02:47Now, throughout most of the course so
  3913. 2:02:49far, what I've done is I've just used
  3914. 2:02:51mostly naive like Opus 4.6 agents in
  3915. 2:02:54order to spawn other Opus 4.6 sub
  3916. 2:02:56agents.
  3917. 2:02:57And that was mostly for capability sake
  3918. 2:02:59because at least in my case cost is not
  3919. 2:03:01a a concern and I really do want to eke
  3920. 2:03:03out the marginal quality um benefits
  3921. 2:03:05wherever possible.
  3922. 2:03:06But there are a lot of cases,
  3923. 2:03:08specifically like enterprise and and big
  3924. 2:03:10infrastructure ones, where, you know,
  3925. 2:03:11people are actually comfortable making a
  3926. 2:03:13minor trade-off for quality. I'm going
  3927. 2:03:16to draw another one of my famous graphs.
  3928. 2:03:18Um people are comfortable making, you
  3929. 2:03:20know, a a a trade-off in terms of, you
  3930. 2:03:22know, cost
  3931. 2:03:24and then quality. Now, in a lot of
  3932. 2:03:26disciplines out there, biology, physics,
  3933. 2:03:28chemistry, and stuff, there's this this
  3934. 2:03:30idea of like this inverted U curve. And
  3935. 2:03:33you can call it whatever the heck you
  3936. 2:03:34want. Um I think the actual name is the
  3937. 2:03:37Yerkes-Dodson curve.
  3938. 2:03:40And what this is is this is basically
  3939. 2:03:41like the the optimal point
  3940. 2:03:45that combines two different factors. And
  3941. 2:03:48so in our case, if we're optimizing for
  3942. 2:03:49both cost and quality, okay,
  3943. 2:03:51simultaneously, not just cost, not just
  3944. 2:03:53quality, you can imagine that the
  3945. 2:03:55optimal place to choose on this graph is
  3946. 2:03:57probably going to be something like over
  3947. 2:03:59here. It's going to be It's going to be
  3948. 2:04:00somewhere around here.
  3949. 2:04:01If you wanted to minimize cost exactly,
  3950. 2:04:03we'd probably go like way over here. But
  3951. 2:04:05obviously we care about quality as well,
  3952. 2:04:06so we're going to push up a little bit,
  3953. 2:04:08right? We don't want this point because
  3954. 2:04:11one, it costs a lot, and then boom, two,
  3955. 2:04:13the quality's pretty low, because
  3956. 2:04:15despite the fact that quality is really,
  3957. 2:04:16really high, um you know, cost is like
  3958. 2:04:18many, many times higher than it was over
  3959. 2:04:20here.
  3960. 2:04:21And so this is sort of like our our
  3961. 2:04:23optimal point. And, you know, a lot of
  3962. 2:04:25large enterprises, since we're dealing
  3963. 2:04:26with hundreds of millions of dollars
  3964. 2:04:28here, um are actually comfortable making
  3965. 2:04:30a little trade-off where, you know, if
  3966. 2:04:31this quality is 85% and then this one
  3967. 2:04:34here is, I don't know, 95%, they're okay
  3968. 2:04:36taking like a like a 10% hit here
  3969. 2:04:40if it means that they also reduce their
  3970. 2:04:42costs by, I don't know, 40% or something
  3971. 2:04:44like that.
  3972. 2:04:46And this is really where all this stuff
  3973. 2:04:47comes in, okay? So, I don't mean to talk
  3974. 2:04:48your ear off here. It's not super
  3975. 2:04:49important, but uh basically uh what a
  3976. 2:04:52lot of people have taken to doing now is
  3977. 2:04:53doing a 60-30-10 rule
  3978. 2:04:55where they'll use some top-level agent
  3979. 2:04:58router, which is sort of like the
  3980. 2:05:00orchestrator in our multi-agent uh
  3981. 2:05:02Chrome window example. And what that
  3982. 2:05:05agent router does is it calls dumber
  3983. 2:05:06models
  3984. 2:05:08and then it assigns different strengths
  3985. 2:05:10to tasks so that, you know, if you're
  3986. 2:05:13giving it a really simple task and you
  3987. 2:05:15say, "Hey, you know, I just want you to
  3988. 2:05:16classify this into one of three
  3989. 2:05:18categories."
  3990. 2:05:19And it's really dumb. It's like, you
  3991. 2:05:21know, red, blue, or green,
  3992. 2:05:23angry, serene, or or healthy, or
  3993. 2:05:26whatever the heck. Um you don't have to
  3994. 2:05:28use like Opus, which is like space-age
  3995. 2:05:30intelligence and costs you a ton more in
  3996. 2:05:32token cost in order to get that done.
  3997. 2:05:35Instead, you know, you can get all the
  3998. 2:05:36way down to like a Haiku or maybe a
  3999. 2:05:38Gemini Flash model or something like
  4000. 2:05:39that. Likewise, if you have some other
  4001. 2:05:41task here, and maybe that task requires
  4002. 2:05:43a lot of um I don't know, research or
  4003. 2:05:45something,
  4004. 2:05:46and you say, "Hey, I want you to go and
  4005. 2:05:47compile 200 million tokens worth of
  4006. 2:05:50stuff and then um give it to me in a big
  4007. 2:05:52report." Well, you know, you probably
  4008. 2:05:53don't want the dumbest model to do that
  4009. 2:05:55for you. But you also don't need the
  4010. 2:05:55most expensive model. So, maybe you'll
  4011. 2:05:57use something like a Saw model or like a
  4012. 2:05:59a lower-level GPT model, which might
  4013. 2:06:00cost two or three million dollars uh two
  4014. 2:06:02or three dollars per million tokens
  4015. 2:06:04instead. And so, you know, this
  4016. 2:06:05allocation where you have 60-30-10, if
  4017. 2:06:08you think about it like sort of a pie
  4018. 2:06:09chart, um what you do is you designate,
  4019. 2:06:12you know, the vast majority of your
  4020. 2:06:15token usage
  4021. 2:06:16to stuff that is in that first category,
  4022. 2:06:18which is kind of, you know, dumber.
  4023. 2:06:21And then what you do is you do the other
  4024. 2:06:2330% or so
  4025. 2:06:25in sort of that mid-tier. And then your
  4026. 2:06:27really, really, really smart models, you
  4027. 2:06:29know, they do the highest-level tasks.
  4028. 2:06:32And basically, what will happen for the
  4029. 2:06:33most part is this would be, you know,
  4030. 2:06:34your Opus 4.6 or your Gemini
  4031. 2:06:393.1 or your GPT 5.4. It'll be
  4032. 2:06:43responsible for routing decisions and
  4033. 2:06:44obviously you want the smartest model
  4034. 2:06:46possible for that. But all of the heavy
  4035. 2:06:47lifting, all the context and stuff like
  4036. 2:06:49that is um through sub-agents that I
  4037. 2:06:51spawned, either Haiku,
  4038. 2:06:53Sonnet,
  4039. 2:06:54or uh you know, I don't know if if if
  4040. 2:06:56you wanted to do a really smart call,
  4041. 2:06:57then you'd obviously spawn an Opus
  4042. 2:06:59sub-agent like me as well.
  4043. 2:07:01And if you do all this, you could
  4044. 2:07:02significantly reduce the cost. I mean,
  4045. 2:07:03like just just think about it
  4046. 2:07:04mathematically. If previously you were
  4047. 2:07:06doing 100 million tokens times, you
  4048. 2:07:09know, $5 per 1 million tokens,
  4049. 2:07:12>> [gasps]
  4050. 2:07:13>> what's the cost there?
  4051. 2:07:14Well, that's obviously going to cost
  4052. 2:07:15$500. And so that's like your Opus only,
  4053. 2:07:18right? But if you did 10 million times
  4054. 2:07:21$5 plus
  4055. 2:07:2430 million times $3
  4056. 2:07:28plus 60 million times
  4057. 2:07:31$1, what's the total cost going to be
  4058. 2:07:33now? Well, it's going to be 60
  4059. 2:07:36plus 90
  4060. 2:07:39plus 50
  4061. 2:07:41or in total, 200.
  4062. 2:07:44And so, you know, 200 expressed as a
  4063. 2:07:45fraction of 500 is 40%
  4064. 2:07:49of our total cost. And we will have just
  4065. 2:07:52saved, you know, 60% with probably
  4066. 2:07:54minimal impacts on quality because the
  4067. 2:07:57things that we're now spawning, you
  4068. 2:07:59know, dumber agents um to do are things
  4069. 2:08:01that, to be honest, the quality was
  4070. 2:08:03already okay a few generations ago back
  4071. 2:08:05with the Haikus and the Sonnets. So just
  4072. 2:08:07to give you an example from something
  4073. 2:08:08that I do pretty often, okay, which is
  4074. 2:08:11going to be some form of lead scraping,
  4075. 2:08:13what you can do is you can actually
  4076. 2:08:14traverse a very large portion of the
  4077. 2:08:17internet using a relatively dumb model,
  4078. 2:08:19these Haiku models. What these do is
  4079. 2:08:22these um scrape a vast majority a vast
  4080. 2:08:25amount of internet data, okay? All of
  4081. 2:08:27the code of uh I don't know, let's say
  4082. 2:08:2910,000 websites or something. And then
  4083. 2:08:32in doing so, they just like that little
  4084. 2:08:33magnifying glass, use some sort of grep
  4085. 2:08:35or extraction prompt to look for things
  4086. 2:08:38that are formatted like email addresses.
  4087. 2:08:40So, if you have like something and then
  4088. 2:08:42it's an at and then it's a you know, the
  4089. 2:08:44term gmail.com,
  4090. 2:08:46odds are this is a real email address,
  4091. 2:08:47right? So, then you store that to a
  4092. 2:08:48database. And because this is just such
  4093. 2:08:50like a mass data application, use HiQ.
  4094. 2:08:52It drives the cost really, really low.
  4095. 2:08:55Well, then maybe um the actual
  4096. 2:08:56enrichment point, you know, takes
  4097. 2:08:58significantly more intelligence. And so,
  4098. 2:09:00maybe here we'll use Sonnet and it'll
  4099. 2:09:01cost us uh $0.008
  4100. 2:09:04per lead, $0.008 per lead. The actual
  4101. 2:09:07outreach part is mostly templated, so we
  4102. 2:09:10use Sonnet for that as well. And then
  4103. 2:09:11maybe at the end we just have a quality
  4104. 2:09:13review step to make sure things aren't
  4105. 2:09:14absolutely nuts. Well, when you do it
  4106. 2:09:16this way, um you know, the math ends up
  4107. 2:09:18being uh $0.008 + $0.005, so that's
  4108. 2:09:21$0.013 + $0.001, $0.014 + $0.015 is
  4109. 2:09:26$0.029.
  4110. 2:09:28And then if you were to go 100% Opus,
  4111. 2:09:30then it would be I don't know, uh uh
  4112. 2:09:32about 12 cents or so per lead. And so,
  4113. 2:09:34on a list of let's just say a thousand,
  4114. 2:09:36which is approximately how much I'm
  4115. 2:09:38sending a day right now. I'm much
  4116. 2:09:39farther down than uh my maximum, but if
  4117. 2:09:42we multiply all these together, we move
  4118. 2:09:43that one, two, three decimal points over
  4119. 2:09:45to the left, then that um ultimately
  4120. 2:09:48would be $15 a day.
  4121. 2:09:51Or, you know, $450 a month. Well,
  4122. 2:09:54instead I'm doing that literally one
  4123. 2:09:57quarter of this. Or something like, you
  4124. 2:09:59know, $120
  4125. 2:10:01a month instead. Obviously, I'd much
  4126. 2:10:03rather the latter. And if my quality is
  4127. 2:10:05only going down a few percentage points
  4128. 2:10:06because of that Yerkes-Dodson curve, you
  4129. 2:10:09know, I'm okay being over here instead
  4130. 2:10:12of over here because this gap to me is
  4131. 2:10:14fine on tasks that aren't super high
  4132. 2:10:16high quality, then uh this is a very,
  4133. 2:10:19very efficient stack. And you know, the
  4134. 2:10:21bigger and bigger my company gets,
  4135. 2:10:23whoever I'm working with, the more the
  4136. 2:10:25cost per lead is going to be important
  4137. 2:10:28versus the actual quality. I've included
  4138. 2:10:30just a little LLM API pricing cheat
  4139. 2:10:33sheet. I don't expect this to be super
  4140. 2:10:35relevant or useful to you guys. There
  4141. 2:10:37are a lot more models for OpenAI
  4142. 2:10:40and Google, but I am actually using um
  4143. 2:10:42not this model series anymore, but this
  4144. 2:10:44one here for some queries. I'm also
  4145. 2:10:46using the flash model series for some
  4146. 2:10:48queries as well.
  4147. 2:10:49And then what's really cool is a lot of
  4148. 2:10:51them offer uh what's called a batch API
  4149. 2:10:53now, where you can submit a bulk number
  4150. 2:10:56of requests over um simultaneously. And
  4151. 2:10:59then if you're comfortable waiting like
  4152. 2:11:01a day or so, what the companies do is
  4153. 2:11:03they batch it and then they serve your
  4154. 2:11:06requests during periods in which they
  4155. 2:11:08have very low inference or low
  4156. 2:11:09competition. So, maybe in the middle of
  4157. 2:11:11the night or something like that. And in
  4158. 2:11:13doing so, they actually get to load
  4159. 2:11:14balance. Like if you think about it like
  4160. 2:11:16if this is like a day
  4161. 2:11:18and then this is their load like on
  4162. 2:11:19their servers and their neural networks
  4163. 2:11:21and stuff, you know, it'll probably peak
  4164. 2:11:22somewhere around like noon and there's
  4165. 2:11:24probably like a couple things and then
  4166. 2:11:25it's like low during the during the day,
  4167. 2:11:27right? I don't know. This is like 4:00
  4168. 2:11:28a.m. What they'll do
  4169. 2:11:31is they'll actually take all of your
  4170. 2:11:32queries, batch them, and then they'll
  4171. 2:11:33just like run them over here when
  4172. 2:11:35there's very little competition.
  4173. 2:11:37And then later when, you know, the
  4174. 2:11:38things go up and stuff like that again,
  4175. 2:11:40um that's okay. And in doing so, what
  4176. 2:11:42they want to do is they want to shift
  4177. 2:11:43some of the the really top end of all of
  4178. 2:11:45these users
  4179. 2:11:47um to the low end to basically fill this
  4180. 2:11:49so that they have a lot more like
  4181. 2:11:50dependable load instead of these like uh
  4182. 2:11:53jagged peaks and whatnot.
  4183. 2:11:55>> [sighs and gasps]
  4184. 2:11:56>> But anyway, don't worry too much about
  4185. 2:11:57that. I just wanted to cover um some LLM
  4186. 2:11:59pricing principles as well so that you
  4187. 2:12:00guys know not only how to manage your
  4188. 2:12:02context better, but also how to um save
  4189. 2:12:06especially when you get into more
  4190. 2:12:07sophisticated multi-agent setups like
  4191. 2:12:09I've been showing you. And that's it.
  4192. 2:12:11Thank you guys very much for watching
  4193. 2:12:12this video end to end. If you guys have
  4194. 2:12:14made it all the way to this point in the
  4195. 2:12:15course, you're part of like the 2 or 3%
  4196. 2:12:18that actually do. Um I'd really
  4197. 2:12:19appreciate a big solid if you could do
  4198. 2:12:21me a favor and subscribe to the channel.
  4199. 2:12:23Something like 70% of you aren't, which
  4200. 2:12:25significantly hurts my reach. And uh
  4201. 2:12:27despite me hating asking for it, it does
  4202. 2:12:29help the channel grow. So, if I've given
  4203. 2:12:30you guys any value whatsoever, please do
  4204. 2:12:32that. You can also send me over a
  4205. 2:12:34comment down below asking any question
  4206. 2:12:36about any point in the video. Um I'm
  4207. 2:12:39much more engaged than the average
  4208. 2:12:40YouTuber, so the probability that I will
  4209. 2:12:41reply is pretty pretty high up there, I
  4210. 2:12:43would say, statistically.
  4211. 2:12:45Um if you guys have any, you know,
  4212. 2:12:47suggestions for future videos or future
  4213. 2:12:48courses, please drop them down below as
  4214. 2:12:50well. And above all else, keep learning
  4215. 2:12:52and growing with AI agents. This is by
  4216. 2:12:55far the biggest and most impactful of
  4217. 2:12:58economic changes that I think any of us
  4218. 2:13:00will see in our lifetime. It's a blessed
  4219. 2:13:02time to be alive in general. You might
  4220. 2:13:03as well not waste it. Make the most out
  4221. 2:13:05of it. All right. Um thank you very
  4222. 2:13:07much. Feel free to use the chapter
  4223. 2:13:08headings to revisit any section in the
  4224. 2:13:10course. And uh looking forward to seeing
  4225. 2:13:11all y'all in the next one. See you
  4226. 2:13:13later.

About this transcript

This page contains the full transcript of AI Agents Full Course 2026: Master Agentic AI (2 Hours) by Nick Saraev, generated from the public captions YouTube serves with the video. The transcript has 28,280 words across 4,226 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.