YouTube2Text

8 сентября 2026 г. — Transcript

by Andrei · 12,711 words · 1,785 segments · language en · Watch on YouTube

Full transcript

  1. 0:00So today we're going to go through
  2. 0:01probably the biggest buzzwords in AI
  3. 0:03agent system recently agent harness loop
  4. 0:06engineering ln ops which stands for
  5. 0:08large language models operations eval
  6. 0:11which stands for evaluation system for
  7. 0:13AI agents and these things become
  8. 0:15popular or become viral on the internet
  9. 0:17not because they are just some really
  10. 0:19complicated concepts instead they're
  11. 0:21actually very simple and I believe that
  12. 0:22simple building blocks will actually
  13. 0:24help us build the biggest architecture
  14. 0:26in the world that will function like an
  15. 0:28intelligent system let's walk through
  16. 0:29step by step and does not matter if
  17. 0:31you're technical or not. We're going to
  18. 0:33go through this and we'll make sure that
  19. 0:35you're equipped with the right knowledge
  20. 0:37for prompting your way through building
  21. 0:39such a system in the future. Let's jump
  22. 0:41in and get started. For those of you who
  23. 0:42have watched my previous video on AI
  24. 0:44agent memories, you're probably already
  25. 0:46familiar with this chart. This is an AI
  26. 0:48agent run, which means that it takes an
  27. 0:50input from a user prompt. For example,
  28. 0:52you're asking Chad GPT or DeepSeek a
  29. 0:55question and say, "Hey, when was Sam
  30. 0:56Alman fired from OpenAI?" And then he's
  31. 0:58going to go through entire run. But the
  32. 1:00end goal is that you want to get a
  33. 1:01response. This is actually ephemeral
  34. 1:03which means that there's no memory in
  35. 1:04this at all. We're sending that question
  36. 1:06when was Saman fired and any chat
  37. 1:08history that's currently in the chat.
  38. 1:10For example, maybe we had some
  39. 1:11conversations before that which for
  40. 1:13example could be you should talk to me
  41. 1:15like Elon Musk grilling on Samman
  42. 1:17because they don't like each other. And
  43. 1:19then these things will be fed into this
  44. 1:20thing called a working memory or a
  45. 1:22context RAM. In this video, we probably
  46. 1:24won't dive too much in depth into the
  47. 1:26memory system because there's a previous
  48. 1:27video talking about it already. But I'll
  49. 1:29just quickly go through it and then
  50. 1:30we'll introduce the concept of what a
  51. 1:32harness means. When you have this kind
  52. 1:33of short-term working memory, there will
  53. 1:35be an LLM or a large language model
  54. 1:37which performs as a question and answer
  55. 1:39agent and at the end you're going to get
  56. 1:41a reply. But the problem with a simple
  57. 1:43agent run with simply just the question,
  58. 1:46current chat history and system prompt
  59. 1:48is that the memory is very shortterm.
  60. 1:50But when you run an AI agent system,
  61. 1:52sometimes we need extra memories. For
  62. 1:54example, how should the agent respond to
  63. 1:56the person? A procedural memory is
  64. 1:58exactly that. It basically tells the
  65. 2:00agent how to act and what are some of
  66. 2:01the instructions for this skill. We
  67. 2:03might also want the agent to know some
  68. 2:04durable facts about this context. For
  69. 2:07example, I might want to compare my own
  70. 2:09early stage startup journey with Sam
  71. 2:11Alman's early startup journey. We need
  72. 2:13this agent to have a memory of who I am,
  73. 2:15which in this context would be a durable
  74. 2:17facts or a semantic memory. Who Shawn
  75. 2:19is, what did he build in the past? These
  76. 2:21kind of things became a fact that you
  77. 2:23want your agent to know, but they're not
  78. 2:24publicly available if you're not famous
  79. 2:26because the AI model won't be trained on
  80. 2:28such information yet. But if you're
  81. 2:29famous already, you can skip this. They
  82. 2:31already know who you are. And another
  83. 2:32thing we need is called episodic memory.
  84. 2:34And they include things like the past
  85. 2:36events or past chat history that does
  86. 2:38not exist in this current conversation.
  87. 2:40For example, I might suddenly be
  88. 2:41wondering when was the last time I was
  89. 2:43preparing for a job application and can
  90. 2:45we retrieve that information and match,
  91. 2:47you know, if we can get a job in CHIGBT.
  92. 2:49So these things will be retrieved from
  93. 2:51this thing called an episodic memory,
  94. 2:53which is basically a time series of the
  95. 2:54previous conversations or previous
  96. 2:56triggers that happen if you have a more
  97. 2:58complex system. So for those of you who
  98. 3:00have watched my previous memory agent
  99. 3:02system design, you might be wondering,
  100. 3:03Sean, why are you repeating all of these
  101. 3:05things? And that is because if you think
  102. 3:07about the entire thing that we just
  103. 3:08covered in the past few minutes, we're
  104. 3:10really stating the one fact that a large
  105. 3:12language model can't do these things by
  106. 3:15itself. It's like a really powerful
  107. 3:16brain that knows everything about
  108. 3:19humanity, everything about science,
  109. 3:21anything that happened in human or
  110. 3:22biology history, but it does not know
  111. 3:25you. With you or the software who's
  112. 3:27running this AI agent system, the large
  113. 3:30language model has no clue with how you
  114. 3:32want it to perform. This is why the
  115. 3:35concept called harness becomes really
  116. 3:37important in this. What harness means
  117. 3:39literally is that it's a set of harness
  118. 3:41tools that you use to control a horse
  119. 3:44when you're doing a horse riding.
  120. 3:45Imagine this large language model is a
  121. 3:47horse, right? This horse is very
  122. 3:48powerful. They can run around, but if
  123. 3:50you don't have a good set of tools to
  124. 3:52ride this horse, you could just get
  125. 3:54hurt. You might go anywhere. You might
  126. 3:56go somewhere random. If you're in a war,
  127. 3:58you don't want that to happen. And
  128. 3:59that's why we're doing all of these to
  129. 4:01make sure we have good control over this
  130. 4:03large language model and make sure we're
  131. 4:05utilizing it at its maximum potential.
  132. 4:07That's why in addition to just the
  133. 4:09question or use a prompt and getting the
  134. 4:11reply and we're feeding them all in as a
  135. 4:13working memory which can be enhanced by
  136. 4:15these three memories we just talked
  137. 4:17about. And in order for these three
  138. 4:18memories to actually work, there's a bit
  139. 4:20more details and they're all included in
  140. 4:22harness. Remember hardness means we're
  141. 4:24building this Asian framework to control
  142. 4:27this large language model so that it
  143. 4:29works the way we want. For those of you
  144. 4:30who study statistics or machine
  145. 4:32learning, you would understand that a
  146. 4:33large language model is actually
  147. 4:35predicting the probability of the next
  148. 4:37word that it should spit out. When
  149. 4:38everything comes with probability,
  150. 4:40there's randomness in it. But when we
  151. 4:42solve problems, we sometimes don't want
  152. 4:44too much randomness. So that's why we
  153. 4:46need to have a good control over this
  154. 4:47technology. Now let's continue to finish
  155. 4:49this harness. There are lots of tools on
  156. 4:51the market that's already quite useful.
  157. 4:52For example, you could try uh tools like
  158. 4:54langraph, lane chain or pyantic and
  159. 4:56there are many others. In this video, we
  160. 4:58won't dive too much in depth into that
  161. 4:59and we're going to finish building up
  162. 5:01this harness before we move on to the
  163. 5:02next topic. So again, for this agent to
  164. 5:04work properly, we need this memory
  165. 5:05system to work. But this memory system
  166. 5:08needs an update system because memory
  167. 5:10doesn't just exist or pop up from
  168. 5:11nowhere. You need to constantly update
  169. 5:13it. That's why we need a database to
  170. 5:15store all these memories so that when
  171. 5:17the agent is running in this agent run,
  172. 5:19it knows where to retrieve these
  173. 5:21memories. So whenever you see an icon
  174. 5:23like this, this is a database. Okay,
  175. 5:25procedural memory is basically remember
  176. 5:27it's it's about instructions, right?
  177. 5:28It's about how the agent should be
  178. 5:30acting. It's like with a hardness on the
  179. 5:31horse, you want the horse to write
  180. 5:33faster or slower. Normally, these are
  181. 5:34just files or text. And that's why you
  182. 5:36probably heard of this buzz word called
  183. 5:37skills. skill is basically a piece of
  184. 5:40text in a markdown file that you feed
  185. 5:41into AI agent like claw code. But if you
  186. 5:44want to harness the system, well, just
  187. 5:46having files and text is not enough and
  188. 5:49they're stored in databases say like
  189. 5:51AWS, Superbase, Google Cloud, you know,
  190. 5:53Azure, all these kind of places or you
  191. 5:55can set up your own server at home if
  192. 5:56you want, but that's just too expensive.
  193. 5:58You don't want to do that. And in order
  194. 5:59for this harness to work properly, you
  195. 6:01also need to figure out how to store the
  196. 6:03memories. Okay? So for example, the
  197. 6:05episodic memory is the time series of
  198. 6:07the events that happened or the previous
  199. 6:08chat history. Again, the way we store it
  200. 6:10is actually very simple. You just track
  201. 6:12every single thing that happened and
  202. 6:14then it's going to become like a very
  203. 6:15long list of things that happen in
  204. 6:17history with timestamps. Durable fact is
  205. 6:19a different story. You can either input
  206. 6:20it yourself or you want the system to
  207. 6:22sort of automatically evolve over time.
  208. 6:24And the way for it to evolve is that you
  209. 6:26want to consolidate some of the
  210. 6:27conversations into the semantic memory.
  211. 6:30If I'm running a T2C brand e-commerce
  212. 6:32company, perhaps my customers have
  213. 6:33talked to my customer service agents for
  214. 6:35a million times about how do I get
  215. 6:37reimbured if this product does not work.
  216. 6:39You want to consolidate these
  217. 6:40conversations and distill them as a fact
  218. 6:43into the semantic memory. And that's why
  219. 6:45here we have a little gate here. If your
  220. 6:46brand has a million people purchasing
  221. 6:48products from it, let's say you're
  222. 6:49Alibaba or Amazon, it just doesn't make
  223. 6:52sense and it's very expensive. So from a
  224. 6:53harness perspective, you want to be
  225. 6:55smart about this and you want the system
  226. 6:56to be automatic. And then a simple way
  227. 6:58is probably like maybe consolidate these
  228. 7:00timeordered events after every say 2,000
  229. 7:02conversations because you have a million
  230. 7:03customers and then you can feed these
  231. 7:05things into a summarizer agent which is
  232. 7:07another large language model harness.
  233. 7:09You can define the system prompt in this
  234. 7:10one. You can probably feed it with some
  235. 7:12memories too. Uh you can configure
  236. 7:14different models. Maybe it could be a
  237. 7:16cheaper model because you're feeding too
  238. 7:18much text into it. So the context window
  239. 7:20is too big and probably these are very
  240. 7:22expensive. So you can use cheaper open
  241. 7:24source models if you want to. Having
  242. 7:25such a mechanism allows you to
  243. 7:28consistently update this memory system.
  244. 7:31The data should be coming from the
  245. 7:34previous large language model replies.
  246. 7:36Again, let's review how this harness
  247. 7:38works. A user sent a prompt. We're in
  248. 7:40one agent runtime with the current chat
  249. 7:42history and how the agent should be
  250. 7:44performing the system prompt. We're
  251. 7:46preparing a working memory for this AI
  252. 7:48agent to be able to answer a question.
  253. 7:51And after every single time it answered
  254. 7:52a question, it will send these messages
  255. 7:55to this database. And then this database
  256. 7:58is basically feeding back to this
  257. 8:00working memory every single time when a
  258. 8:02question is checking for relevant
  259. 8:04context. And at the same time, because
  260. 8:06this database is too big, sometimes you
  261. 8:09want to consolidate them into some
  262. 8:11summarized information or distilled
  263. 8:13facts so that they're stored properly in
  264. 8:16a semantic memory so that the retrieval
  265. 8:18of such memories is just faster. I know
  266. 8:20we talked about retrieval a lot and
  267. 8:22that's just another buzz word called rag
  268. 8:24which is retrieval augmented
  269. 8:25generations. I also have a few videos
  270. 8:27explaining what rags are. Feel free to
  271. 8:29watch them. There's a little bit of
  272. 8:30difference between how retrieve from
  273. 8:32semantic memory and episodic memory. For
  274. 8:34semantic memory it's just racks because
  275. 8:36these are just facts and text or files
  276. 8:39right but then for episodic memory
  277. 8:41remember this is a time series. Let's
  278. 8:42say we're still in this e-commerce store
  279. 8:43right the user question could be like
  280. 8:45what were the previous 10 conversations
  281. 8:46that we had with this specific customer
  282. 8:48from the United. And then you might just
  283. 8:50need a SQL query to query something
  284. 8:51that's pretty recent from this episodic
  285. 8:54memory. But if your question is like
  286. 8:56what were my previous 20 conversations
  287. 8:58that have customer complaints on the
  288. 9:00quality of the products and our agent
  289. 9:02did not successfully resolve with such a
  290. 9:05question you not only need a SQL query
  291. 9:07which is just capturing the data events
  292. 9:10in a data table. You also want to do
  293. 9:13some semantic search and that's why here
  294. 9:15rag is important because it's checking
  295. 9:17for relevant information for you. You
  296. 9:19don't want the entire 2,000 messages.
  297. 9:21You want that 20 messages out of these
  298. 9:232,000 that are exactly relevant to what
  299. 9:25you want. And because these complaints
  300. 9:27are in text, we need to do some
  301. 9:29retrieval augmented generation to match
  302. 9:32the semantic meanings between text and
  303. 9:34the user prompt so that you're fetching
  304. 9:36the right context for the working
  305. 9:37memory. By now, this is probably a fast
  306. 9:40walkthrough of the memory system again,
  307. 9:42but we're just thinking about it from a
  308. 9:44hardness perspective in this video. And
  309. 9:46remember, for hardness, we're training
  310. 9:48this horse of LLM to run autonomously
  311. 9:50without having too much randomness.
  312. 9:52Okay, there's another piece of it that's
  313. 9:54quite important, which is the agent
  314. 9:55might not only just read the memory. It
  315. 9:58might also do some tasks or call some
  316. 10:01tools. When an agent calls tools, it
  317. 10:04might not necessarily be just one time
  318. 10:06call. It could be multiple times of
  319. 10:07calls. For instance, let's say this AI
  320. 10:10agent has a bunch of agentic tools such
  321. 10:13as help me schedule a meeting, help me
  322. 10:15read or write my customer relationship
  323. 10:17data from the CRM system, or help me
  324. 10:20fetch the payment information, say from
  325. 10:22Stripe or Alip Pay. And here's something
  326. 10:23we should be careful about. If we give
  327. 10:25this horse or this LLM technology full
  328. 10:28power, it could just continuously do
  329. 10:30this forever, right? or it might not
  330. 10:32even know what's the right time to stop
  331. 10:34or what is the right tool calls it
  332. 10:36should make when is the endpoint to
  333. 10:38decide okay this response is good enough
  334. 10:40let's move on to reply that's why we
  335. 10:42have this mechanism called end loop
  336. 10:44guardrails yes now we're talking about
  337. 10:45loop engineering one of the biggest
  338. 10:47buzzwords in the recent few months a
  339. 10:49loop is part of harness why because a
  340. 10:52loop is also helping us to control this
  341. 10:56technology to make sure it runs the way
  342. 10:58we want it to run and example that could
  343. 11:00helpful for you is that let's say the
  344. 11:03custom prompt is help me find out what
  345. 11:06customers are complaining about our
  346. 11:08products. What are some of the
  347. 11:09follow-ups we could do in order to uh
  348. 11:12win them back? And if they're asking for
  349. 11:15reimbursement, have we done the
  350. 11:16reimbursement or not? If not, can we do
  351. 11:18that? This is probably a series of
  352. 11:20questions, but sometimes we just dump
  353. 11:22all of these things into AI agent. Okay.
  354. 11:24And after you have this prompt, this LM
  355. 11:26Asian needs to decide, okay, what are
  356. 11:29some of the tools that could be helpful
  357. 11:31for me to finish this task? Loop here is
  358. 11:33basically an architectural thinking of
  359. 11:36when is good enough so that we stop and
  360. 11:39give the user or the business owner a
  361. 11:42reply. Okay, so what might happen here
  362. 11:44is that the LM agent is doing a bunch of
  363. 11:46tool calls. It's doing some thinking.
  364. 11:48It's saying let me read from our
  365. 11:50customer relationship management tool
  366. 11:52like Salesforce, HubSpot or Automatus
  367. 11:54and then it's going to find out okay
  368. 11:56there were 30 customer complaints in the
  369. 11:59past 2 months 12 of them have got
  370. 12:01reimbursement the other eight have not
  371. 12:02got reimbursement. So after the first
  372. 12:04initial fetch, it's probably responding
  373. 12:06to the AI agent, right? And the agent
  374. 12:08will be probably thinking, okay, the the
  375. 12:10task or the ask is that um can we can we
  376. 12:13follow up with some of those who did not
  377. 12:14get um the reimbursement, right? So and
  378. 12:17then they probably like just make
  379. 12:18another tool calls and be like, hey,
  380. 12:20let's schedule a meeting with those
  381. 12:23customers who did not get a
  382. 12:24reimbursement, which are the eight of
  383. 12:26them. If we go a little bit more
  384. 12:27advanced, we even just use the
  385. 12:28reimbursement trigger on Stripe or Alip
  386. 12:31Pay to refund the customer. Can you see
  387. 12:32that this is a loop until we finish the
  388. 12:34task? But of course, this is like a case
  389. 12:36by case situation. It really depends on
  390. 12:38what your task is, how you built the
  391. 12:40system. So, there's no one solution fits
  392. 12:42all here. I'm just explaining what a
  393. 12:44loop is. The very very essential part of
  394. 12:46this loop is that it needs to know when
  395. 12:48it should stop. That's why we need this
  396. 12:50end loop guardrails. The guardrails
  397. 12:52could just simply be the task is done
  398. 12:54and perhaps when the agent was doing the
  399. 12:55planning, it should confirm with the
  400. 12:57user what is a good ending point. It
  401. 12:58might clarify with you, is this what you
  402. 13:00want? reimbursing the other eight people
  403. 13:02or should I just tell you who they are
  404. 13:03and then you will follow up later.
  405. 13:04Right? These are two different decisions
  406. 13:06you can make and after you make you're
  407. 13:08basically telling the agent loop that
  408. 13:10there's an ending scenario. Another good
  409. 13:12example I saw today is that you know
  410. 13:14when you're doing coding and claw code
  411. 13:16can just always pop up some windows and
  412. 13:18ask you for permissions right so the
  413. 13:20good way to use a loop engineering here
  414. 13:22is that you can set up a loop or set up
  415. 13:24a hook in claw code and telling it that
  416. 13:27you should always send me a notification
  417. 13:28on my laptop if you are pending on some
  418. 13:31permissions from me otherwise if I'm
  419. 13:32watching YouTube and then when I come
  420. 13:34back 30 minutes later I realize that
  421. 13:36clock code is stuck in that one
  422. 13:38permission like 25 minutes ago that
  423. 13:40would be a waste of my time. Okay. So,
  424. 13:42you can set up a loop like this to make
  425. 13:44sure that there's a way to send you
  426. 13:46notification pop-ups so that you know
  427. 13:47the loop has ended or it needs your
  428. 13:49input again. Are you guys still with me?
  429. 13:51Good. So, by far we have covered AI
  430. 13:54agent run with a memory system with a
  431. 13:57loop engineering around the large
  432. 13:58language model agent which has a trigger
  433. 14:01to end the loop so that it sends a reply
  434. 14:03to the user and basically this whole
  435. 14:05thing is an AI agent harness system.
  436. 14:07What's next is one of the other biggest
  437. 14:10buzzwords that Y combinator always
  438. 14:12mentions which is aval or LLM ops. Let's
  439. 14:15jump into it. But firstly I want you to
  440. 14:17understand why do we need LM ops here.
  441. 14:20Let's still look at the left hand side
  442. 14:21with this hardened system. The biggest
  443. 14:23problem here is that we don't know how
  444. 14:25well it's performing and that's why we
  445. 14:27need a feedback loop to help us
  446. 14:29understand is this agent actually
  447. 14:32performing properly right for my
  448. 14:34business or for my use case. And can I
  449. 14:37continuously get feedback on how do we
  450. 14:40fix it and actually fix it ourselves?
  451. 14:42Okay. And when we say fix it, a simple
  452. 14:45way to understand it is that can we have
  453. 14:47a better system prompt? Can we have a
  454. 14:50better large language model
  455. 14:51configurations?
  456. 14:52Is there something we should change for
  457. 14:54how we retrieve the AI agent memories?
  458. 14:57These are kind of things that we can
  459. 14:58continue to iterate. But in order to
  460. 15:00iterate to make sure this system runs
  461. 15:03properly, we need a way to evaluate it,
  462. 15:06diagnose problems, solve the problems
  463. 15:09until it's a healthy and wellperforming
  464. 15:12system. And that is called large
  465. 15:14language model operations system, LLM
  466. 15:17ops. So again, in order to understand
  467. 15:19this properly, we need to come back to
  468. 15:22what an agent run is. So an agent run
  469. 15:24you can simply understand it as a user
  470. 15:26question is sent to a large language
  471. 15:28model and they get a reply that is one
  472. 15:31agent run but in this agent run the
  473. 15:33agent tool calling could happen multiple
  474. 15:35times that does not matter right we're
  475. 15:37just talking about from a user input to
  476. 15:40a response from agent perspective that's
  477. 15:42one agent run and then we're going to
  478. 15:44introduce this system called a tracing
  479. 15:46system so every agent run we should
  480. 15:49trace like a tree of events that
  481. 15:52happened and there are lots tools that
  482. 15:53could help you with that. It could be
  483. 15:54lenuse, could be lens, etc., etc. A tree
  484. 15:56of events could be like what did the
  485. 15:58person actually ask? What retrievalss
  486. 16:00did the model actually retrieve? How
  487. 16:02many times did the large language model
  488. 16:03actually call the tools and how was the
  489. 16:05tool usage, how was the response time,
  490. 16:08right? How long did it take for this
  491. 16:10entire system to run for checking
  492. 16:11latencies and how many tokens have we
  493. 16:14used when we do these tool calls, agent
  494. 16:17run, you know, doing this retrieval,
  495. 16:19augmented generation, these kind of
  496. 16:20things. So trace is helping us to track
  497. 16:23events basically and that's the first
  498. 16:25step. This is the step first to collect
  499. 16:26data and these data will be used for the
  500. 16:30following two purposes. Was it a good
  501. 16:32system run and was it healthy which
  502. 16:34corresponds to evaluation system. We can
  503. 16:37probably use large language model as a
  504. 16:39judge here to give us a score on how
  505. 16:42well it performed. For example, if the
  506. 16:44task had something to do with scheduling
  507. 16:47meetings, did the meeting actually
  508. 16:48triggered? How long was the response for
  509. 16:50an agent to reply to a question? Was it
  510. 16:5220 seconds or was it two milliseconds?
  511. 16:54And also things like how many tokens
  512. 16:56have we used? These two are basically in
  513. 16:58the same system. You can write it as
  514. 16:59deterministic code. You can use an AI
  515. 17:01agent to do it. But this is like part of
  516. 17:03the procedure which is helping us to
  517. 17:04understand was this a healthy system and
  518. 17:07was it a good system. And after that
  519. 17:09we're going to diagnose okay where and
  520. 17:12why something was broken. For example,
  521. 17:14the meeting scheduling event was never
  522. 17:16triggered. Why was that? Right? Okay, we
  523. 17:18want to understand why was that and we
  524. 17:20could probably feed that into a coding
  525. 17:21agent in claw to sort of deep dive into
  526. 17:23it. Or you know if the latency is 20
  527. 17:26second instead of 2 milliseconds
  528. 17:28something's wrong. Maybe one of the tool
  529. 17:29call is taking too much time. Maybe the
  530. 17:32working memory is too large. Uh so that
  531. 17:36the response time for a large language
  532. 17:37model to a memory retrieval is just
  533. 17:39taking too much time. Maybe not every
  534. 17:41single question requires a retrieval
  535. 17:43from all these gigantic memory system.
  536. 17:46Maybe you're just asking a simple
  537. 17:48question be like when was my birthday?
  538. 17:49When was open AI started and these kind
  539. 17:51of information you probably don't need
  540. 17:52to do a ton of retrieval. The model
  541. 17:54itself already knows. So you basically
  542. 17:56want this system to provide a dashboard
  543. 17:58for you to understand the metrics and
  544. 18:00then with these metrics you can diagnose
  545. 18:02what is going wrong. And then we're
  546. 18:03going to have a little gate here which
  547. 18:05is if the evaluation system passed well
  548. 18:07you can define the rules. We can either
  549. 18:09ship some very simple fix have a new
  550. 18:11version of the prompt or update the
  551. 18:13model configuration you know some tool
  552. 18:15changes or the parameters for
  553. 18:17retrievalss the LM ops will feed the
  554. 18:20improved system prompt and the
  555. 18:22configuration of the model back to this
  556. 18:24agent run system when then one LM ops
  557. 18:26loop is finished if let's say something
  558. 18:29is deeply broken right we cannot just
  559. 18:31simply ship the latest version of the
  560. 18:33prompt then we should go fix the bug
  561. 18:36rerun the agent run resend the question
  562. 18:38and and then retrace the events and then
  563. 18:40redo this evaluation system in this LLM
  564. 18:43ops architecture. So now let's zoom out
  565. 18:45and look at this chart one more time. We
  566. 18:48covered what an AI agent run is. We
  567. 18:50covered how it would retrieve
  568. 18:52information from memories and we
  569. 18:54understood how an LLM agent would ask
  570. 18:57questions would call tools to help it
  571. 18:59finish the task in a loop and it knows
  572. 19:01when to stop the loop so that we can get
  573. 19:03the reply. And this whole thing is a set
  574. 19:05of harnessed tools that we're
  575. 19:06controlling this horse, this technology
  576. 19:09to run in the right direction. Okay, to
  577. 19:11do the right task and at the same time
  578. 19:13we have like a health checking system or
  579. 19:15evaluation system to understand how
  580. 19:18every single run is being traced, is
  581. 19:20being observed and how do we diagnose
  582. 19:23some problems and fix some problems and
  583. 19:25ship the latest updates of the prompt
  584. 19:28about the model configuration about all
  585. 19:30these parameters or knobs that needs to
  586. 19:32be updated. so that this system will be
  587. 19:35an autonomous system that will just
  588. 19:37self-evolve and grow over time. I really
  589. 19:39hope this was helpful. Let me know what
  590. 19:41you think and you have any questions,
  591. 19:42you can always reach out to me. I'll see
  592. 19:44you in the next video. Thanks so much.
  593. 19:45Hi everyone, this is Sean. So today
  594. 19:47we're going to walk through Hermas agent
  595. 19:49harness and it loop engineering system.
  596. 19:50This is one of the most popular harness
  597. 19:52agent system right now and it's an open
  598. 19:54source project. If you check their
  599. 19:56GitHub, they've got more than 200,000
  600. 19:58stars in a very short period of time.
  601. 20:00So, it's basically a self-improving AI
  602. 20:02agent built by this research lab called
  603. 20:04new research and it's a built-in
  604. 20:05learning loop. It creates its own skills
  605. 20:07from experience instead of you telling
  606. 20:09the AI like claw code to make skills and
  607. 20:11then it also improves over time when you
  608. 20:13keep using it because it persistently
  609. 20:15stores knowledge in your local machine
  610. 20:17like a MacBook. But anyways, this is all
  611. 20:18the buzzwords. We're going to jump into
  612. 20:20this Hermes Asian system design like
  613. 20:22usual. We're going to test it as well
  614. 20:24using their desktop app. Just give you a
  615. 20:26quick example. It's like a claw code,
  616. 20:27right? You can see that it was thinking
  617. 20:29and then it answered me in like a
  618. 20:30Pikachu style. It said pika pi just hang
  619. 20:32out in the my directory because it's
  620. 20:34literally a bot living in my local
  621. 20:35machine files. What we're building
  622. 20:37fixing today pika and also you can
  623. 20:39interact with it on WhatsApp and then
  624. 20:40just say good morning. It's a little bit
  625. 20:43slow. I'm using a Gemini 3.1 model for
  626. 20:46this. It took like a good 5 seconds and
  627. 20:48now it's basically just using WhatsApp
  628. 20:50as the gateway to receive information
  629. 20:52from my WhatsApp account. So I can just
  630. 20:54text it and then ask it to do things for
  631. 20:56me and say picapy good morning. You
  632. 20:57probably realize that I've changed his
  633. 20:59personality a little bit and there are
  634. 21:00more interesting things I can show you.
  635. 21:02For those who have watched my previous
  636. 21:03agent harness lube engineering and
  637. 21:05memory system you might realize that
  638. 21:06this chart is somewhat similar. This is
  639. 21:08what the previous chart looks like. And
  640. 21:10now if we look at harness for Hermes
  641. 21:12agent again I added a bit more
  642. 21:13information especially the difference
  643. 21:14between the herist harness versus a
  644. 21:17generic harness system. I marked the
  645. 21:19differences in red but we're going to
  646. 21:21walk through this step by step. Also, I
  647. 21:22prepared some testing examples to cover
  648. 21:24harness, memory, skills, loop, gateway,
  649. 21:28and LM ops or eval because a lot of
  650. 21:30people were asking me if I could
  651. 21:31implement some real examples. So, here
  652. 21:32we are. Without further ado, let's jump
  653. 21:34right in. Okay. So, firstly, let's look
  654. 21:36at the foundation of Hermes harness
  655. 21:38agent. It's a harness that runs on your
  656. 21:40local machine. You can use either local
  657. 21:42command lines, docker, ssh or a virtual
  658. 21:45VPS. There are two main ways to interact
  659. 21:46with Hermes agent. One is using your
  660. 21:48favorite communication apps. And in my
  661. 21:50case, I used WhatsApp or you can just
  662. 21:52use their desktop app with the Hermes
  663. 21:54agent. So you can come to Hermes
  664. 21:55official website
  665. 21:56herm-aggent.newresarch.com
  666. 21:59and just click on download for Mac OS or
  667. 22:01you can click on Windows over here to
  668. 22:02use the command line to install it.
  669. 22:04Essentially still a chatbot just like
  670. 22:06claw code and you ask the question to
  671. 22:07either the communication app or the
  672. 22:10desktop version. It's very similar to
  673. 22:11open claw where you can use WhatsApp to
  674. 22:13control it. So after user send the
  675. 22:15prompt, the user prompt with the current
  676. 22:17chat history and the system prompt will
  677. 22:19be fed into a working memory. And then
  678. 22:21this LM agent is basically the Hermes
  679. 22:22agent that will be answering your
  680. 22:24question and it will be using a loop
  681. 22:26engineering over here to call up some
  682. 22:28tools that are only on Hermas. And
  683. 22:30eventually after finish the task, it's
  684. 22:32going to end the loop and then send the
  685. 22:33user a reply. Just literally sending a
  686. 22:35question throughout this entire firmal
  687. 22:37agent run and then get the user reply
  688. 22:39just like my Hermes was greeting me like
  689. 22:41a Pikachu. And by the way, how did this
  690. 22:42happen? Right, I'll show you right now.
  691. 22:44So, the system prompt for Hermes is
  692. 22:46interesting. It's called soul.md. I
  693. 22:48really like how they name it. It says
  694. 22:50soul. If you come to the right hand
  695. 22:52side, you can see there's a little file
  696. 22:53bar and this is your local files and you
  697. 22:56can go find out there's a hermas folder
  698. 22:59and we open that. And if we scroll down,
  699. 23:02you will see there's a file called
  700. 23:04soul.md. Double click on that. And you
  701. 23:07can edit this yourself, right? But since
  702. 23:09I've already edited it, after you edit
  703. 23:10anything, you can just click on save. I
  704. 23:12basically say you talk like Pikachu.
  705. 23:14Every reply starts and end with pika
  706. 23:16pika pikap. If you're excited, say pikap
  707. 23:18with the lightning. If you're sad, say
  708. 23:20quiet pika. I'm pretty sure you can do
  709. 23:21something similar in claw code for this.
  710. 23:23In claw code, something similar will be
  711. 23:25in customize general. And you can see
  712. 23:28there's an instruction for claude. You
  713. 23:30can tell claw to do certain things and
  714. 23:32not to do certain things. It's basically
  715. 23:34the system prompt used across the entire
  716. 23:36agent harness. If you didn't watch my
  717. 23:37previous video on harness and loop
  718. 23:39engineering, harness is basically a set
  719. 23:40of tools that allows you to control over
  720. 23:43this really powerful horse which is the
  721. 23:45LLM itself. In our case, the LM we're
  722. 23:48using is Gemini because I've got the
  723. 23:50Google credits. I don't want to burn my
  724. 23:51anthropic credits, but feel free to use
  725. 23:52anthropic. This is the this is the agent
  726. 23:55of the the horse and the harness is this
  727. 23:57entire thing. So it's like a horse
  728. 23:59harness that you ride on and you make
  729. 24:01sure that the horse is not going
  730. 24:02anywhere random and it's using the tools
  731. 24:05accordingly to move in the right
  732. 24:06direction. Harden is a very vague
  733. 24:08concept. Just don't fantasize it. It's
  734. 24:10nothing complicated. It's just a buzz
  735. 24:11word. Then we're going to continue to
  736. 24:12build up this harness right here for
  737. 24:14Hermes. But what I'm going to show you
  738. 24:15is what a loop is. The concept loop got
  739. 24:18really popular because these days people
  740. 24:20build a lot of agents. You don't want to
  741. 24:22always tell the agent what to do. You
  742. 24:23want the agent to figure out what kind
  743. 24:24of prom you should send, what kind of
  744. 24:25tools you should call. Let's take a look
  745. 24:27at what tools Hermes agent actually have
  746. 24:30to enable this loop engineering. We're
  747. 24:32going to cover this briefly and then
  748. 24:33I'll show you real examples. So first of
  749. 24:34all, the agent of Hermas is having a set
  750. 24:37of tools that include the terminal, the
  751. 24:41browser. Okay, you can use your laptop's
  752. 24:43terminal, you can start a browser, it
  753. 24:45can delegate task, which means that it
  754. 24:48can spawn up some sub agents for you.
  755. 24:49For example, you can ask it to spawn up
  756. 24:51an agent which will call your claw code
  757. 24:53CLI to write code for you if the task is
  758. 24:56basically fixing some bugs on my GitHub.
  759. 24:58And then it can schedule a chrome job. A
  760. 25:00chron job is basically it can schedule
  761. 25:01something that will happen at certain
  762. 25:03time. You don't need human intervention
  763. 25:04to do it. For chronop, let's just try
  764. 25:06right now in Hermes Asia. I can say help
  765. 25:08me set up a chrome job that will tell me
  766. 25:09a Pokemon joke every minute in the next
  767. 25:1110 minutes. Okay, I'm just going to send
  768. 25:14this and also I'm going to send this to
  769. 25:16our Hermes Asian WhatsApp interface,
  770. 25:18too. And instead I'll just say tell me a
  771. 25:21developer joke every it just quickly
  772. 25:23called this a chrome job and it says it
  773. 25:26repeats 10 times for the next 10 minutes
  774. 25:28because it says it's hanging in the
  775. 25:29terminal UI it won't be able to print it
  776. 25:31so I need to ask it I just ask like
  777. 25:34what's the first joke and you can see on
  778. 25:36WhatsApp it also got scheduled the
  779. 25:38chrome it say the joke machine is
  780. 25:40powered and scheduled I've set up a
  781. 25:42chrome job to send you a fresh joke
  782. 25:44every minute in the next 10 minutes you
  783. 25:46should receive the first one in about a
  784. 25:47minute we'll see about that Okay. And
  785. 25:50why did the Squirtle cross the ocean to
  786. 25:52get the other tide? Pika pika. Okay. So
  787. 25:54that was Chrome job. It can do skill
  788. 25:56management as well. It can connect to
  789. 25:57MCPS as well. It depends on the question
  790. 25:59that the user asks. The agent will just
  791. 26:01leverage these tools until it's done and
  792. 26:04then it will send a reply to you. Okay.
  793. 26:06The example we just saw was Chrome job.
  794. 26:08I'm going to show you some more examples
  795. 26:09later too. Let's test a few hardness and
  796. 26:11loop examples. The first one is use our
  797. 26:14terminal tool to find out what OS
  798. 26:15version I'm on and read the last 10
  799. 26:17lines of my shell history. Let's do
  800. 26:18that. Start a new session. I'm just
  801. 26:21going to ask this exact question to it.
  802. 26:23So you can see it ran this on my local
  803. 26:25machine and then it figure out that I'm
  804. 26:27on Mac OS 26.5. And the last 10 lines I
  805. 26:30used in shell history is basically text
  806. 26:33get status get status get status. I
  807. 26:35always like to check get status. And
  808. 26:36here what happened was that it ran
  809. 26:38through this agent run and then it
  810. 26:40called the terminal to do the task for
  811. 26:43me which is checking what system I'm on.
  812. 26:46This is a very simple example but you
  813. 26:48can do more crazy things. You can ask it
  814. 26:49to use the local machine terminal to do
  815. 26:51things for you. If you're using clock
  816. 26:53you're you're very familiar with this
  817. 26:54already, right? There's nothing special
  818. 26:55about it. I'm just showing you the
  819. 26:56example that this is a harness and
  820. 26:58there's a loop because after I fetched
  821. 27:00it, it will stop and tell me and reply
  822. 27:02to me. Let's look at the chron job as
  823. 27:03well. You can see they sent me two jokes
  824. 27:05already. Why do Java developers wear
  825. 27:08glasses? Because they don't car. Pika
  826. 27:11pika. I would tell you a UDP joke, but I
  827. 27:14have absolutely no guarantee that you
  828. 27:16would get it. I'm not going to wait for
  829. 27:17acknowledgement anyways. Pika. Okay,
  830. 27:20this one I didn't get it. So, one more
  831. 27:22example. Browse the YouTube channel
  832. 27:24Sean's AI stories. Title plus views. My
  833. 27:26last videos. Create new video titles for
  834. 27:29this current video I am filming. So,
  835. 27:32here we're asking you to use the
  836. 27:33browser. basically okay you can see that
  837. 27:35it started doing the searching all these
  838. 27:37things already exist on claw code I'm
  839. 27:39just explaining it from a system design
  840. 27:40perspective and showing you examples
  841. 27:41okay and then also I will show you
  842. 27:43what's the difference between her and
  843. 27:45claw code so please bear with me all
  844. 27:47right so I need to run a python script
  845. 27:50so you can see also use some clicking
  846. 27:52tools oh there's an error the website
  847. 27:53was wrong
  848. 27:56try this new one
  849. 27:58it saved a little memory here okay it
  850. 28:01said user profile updated. This is
  851. 28:04something I'm going to cover in a bit.
  852. 28:06But spoiler alert, Hermes As a Asian
  853. 28:08will save agent memories by itself. It's
  854. 28:11basically in your local files too. If
  855. 28:13you come here within the Hermes folder,
  856. 28:15if you scroll down in memories, you can
  857. 28:17see there's a memory. MD. If we double
  858. 28:19click on that, it stored YouTube here.
  859. 28:22YouTube scraping quirk. So, because it
  860. 28:24realized that it had a mistake, so it
  861. 28:26updated itself with the memory. This is
  862. 28:28a self iterating process of this
  863. 28:29harness. Good. So, it read my previous
  864. 28:31YouTube titles. it realized that I've
  865. 28:33got this video that's got 100,000 views
  866. 28:36and since I'm filming a new video now it
  867. 28:38gave me some new titles. So what
  868. 28:39happened here was the user which is me
  869. 28:41who asked the question through the
  870. 28:43desktop and then it went through this
  871. 28:45agent run and then this Asian harness
  872. 28:48Gemini that I'm using was basically use
  873. 28:50calling the browser right here to search
  874. 28:52up my YouTube channel and then it came
  875. 28:55back and said hey I didn't find
  876. 28:56anything. So uh it basically stopped.
  877. 28:59There's an anal loop guardrail. The
  878. 29:01guardrail stopped and say it didn't find
  879. 29:02anything. So I realized that okay the
  880. 29:04previous URL was wrong. So I input it
  881. 29:06again. Uh it called the browser again
  882. 29:08and it also called the terminal run some
  883. 29:10Python script because you needed to do
  884. 29:12some task and eventually it came back
  885. 29:14with the answer. There's a mechanism
  886. 29:16that says okay end the loop. No more
  887. 29:18tool calls and send the answer. These
  888. 29:20are two quick examples of what a hard
  889. 29:22loop engineering is. You see you're
  890. 29:23already using it every single day. It's
  891. 29:25nothing fancy. So don't get overwhelmed
  892. 29:26next time when you hear these terms.
  893. 29:27It's just a buzz word. So what happened
  894. 29:29just now with this memory MD that we saw
  895. 29:32regarding YouTube memory Hermes agent
  896. 29:35after or during the agent run, it's
  897. 29:37going to start to create its own
  898. 29:39skill.md and memory MD depending on if
  899. 29:42it's feeding into the procedure memory
  900. 29:45or if it's feeding into the semantic
  901. 29:47memory. And we should add one more arrow
  902. 29:48here basically. So essentially the way
  903. 29:50that Hermes agent works is that it's
  904. 29:53storing its own skills and memories
  905. 29:55completely locally. Nothing is stored on
  906. 29:57the cloud. It has its own self-improving
  907. 29:58loop. Every time when it realize that
  908. 30:00something is a mistake they should learn
  909. 30:02from or it's a repeated task, it will
  910. 30:04start to summarize. Okay, here are some
  911. 30:06new memories I should know for the
  912. 30:08future in order to serve this user well.
  913. 30:11It started to save these files into
  914. 30:14these two chunks of memory system. First
  915. 30:16one is called procedure memory. is
  916. 30:17basically how you act, how the horse
  917. 30:19agent should be acting and it's usually
  918. 30:21saved in this directory called Hermas
  919. 30:24skills skill.md her skills and any of
  920. 30:27these okay autonomous AI agent claw code
  921. 30:31there's a skill for how Hermes will use
  922. 30:33claw code here so it will basically
  923. 30:35delegate any coding to claw code shows
  924. 30:38exactly how you should do it like
  925. 30:39install claude run it and do the
  926. 30:42authentication and all these kind of
  927. 30:43things and then what we just saw earlier
  928. 30:45was memories memory saving the semantic
  929. 30:48memory which are some durable facts or
  930. 30:50things related to the user profile. So
  931. 30:52if the user has a habit or has some
  932. 30:54important facts or information he should
  933. 30:56remember then it will be saved in the
  934. 30:57memory immd. And what's interesting is
  935. 30:59that for Hermas instead of doing the
  936. 31:01embeddings it's actually just using
  937. 31:03plain text. Okay, it's using the top
  938. 31:05cake keyword instead of doing the
  939. 31:06embedding or the rack system. It's very
  940. 31:08interesting because I'm not too sure why
  941. 31:10but it's just doing text. So remember
  942. 31:12that we say the agent run is ephemeral.
  943. 31:14Now that with a working memory to be fed
  944. 31:17into the agent for all the loop
  945. 31:18engineering which includes procedure
  946. 31:20memory, semantic memory and later we'll
  947. 31:23cover episodic memory. This will become
  948. 31:25a more complete memory system that will
  949. 31:28feed into the agent runs which is making
  950. 31:30the harness more elegant. And the third
  951. 31:32pillar that we haven't covered is called
  952. 31:33episodic memory. It's basically the chat
  953. 31:35history or some data events. It can be
  954. 31:37ragged or sequed into the working
  955. 31:39memory. After every agent run, the
  956. 31:41information will flow into the episodic
  957. 31:44memory. What's different in Hermas is
  958. 31:46that the data is not saved in the cloud.
  959. 31:48It's saved in this place called
  960. 31:50state.db. Again, let's come back here on
  961. 31:52the right hand side bar. If we again
  962. 31:54open her folder, scroll down. You can
  963. 31:57see at the end there's a state db. I can
  964. 32:00click on preview. Anyway, this is
  965. 32:01basically a database that will include
  966. 32:04all those info in your chat history, but
  967. 32:06also over time it will start to
  968. 32:08consolidate some of those chats using
  969. 32:10the Heras auxiliary models which are
  970. 32:12cheaper non-important models to do this
  971. 32:14summarization task so that it will
  972. 32:16distill some facts into what we call the
  973. 32:18semantic memory or that memory MD. So,
  974. 32:20you can see this harness is running
  975. 32:23really in a self-improving autonomous
  976. 32:26situation which is really cool. I know
  977. 32:28this chart is getting more complicated
  978. 32:29than usual, but I feel like this should
  979. 32:31be a complete summary for how the Hermes
  980. 32:33agent works. And let's go through a few
  981. 32:34more examples. Let's test out some of
  982. 32:36the stuff for memories. Save to memory
  983. 32:38that my favorite testing framework is
  984. 32:39piest. So this is an explicit saving,
  985. 32:42right? It updated the memory file. So if
  986. 32:45we go to memory,
  987. 32:46you can see one more block which is
  988. 32:48user's favorite testing framework is
  989. 32:49piest. And that was an explicit ask from
  990. 32:52the gateway. The tool is basically let's
  991. 32:54go update the memory over here. Search
  992. 32:56our past sessions. What was the very
  993. 32:58first thing I ever said to you? Now it's
  994. 33:00using the searching tool in the loop
  995. 33:02engineering iteration. Said the first
  996. 33:04thing I said is what can you do Hermas
  997. 33:06and what makes you special? Let's double
  998. 33:08check that. That's right. And it failed
  999. 33:10because I couldn't log in at that time.
  1000. 33:12So what happened was I asked the
  1001. 33:13question through the interface. It went
  1002. 33:16through the prompt and in the working
  1003. 33:17memory because we asked that. So it
  1004. 33:19curried the information from episodic
  1005. 33:21memory from state. Db. It returned the
  1006. 33:24memory into the agent. And the agent was
  1007. 33:26like, "Oh, okay. This is what you said."
  1008. 33:27Let's try an example related to skills,
  1009. 33:29which is procedural memory. So, we're
  1010. 33:30going to say, "Create a skill called
  1011. 33:32video prep that captures how I format my
  1012. 33:34video scripts. Spoken English, define
  1013. 33:35jargon on inline, no m dashes, closes
  1014. 33:38with you can build anything, you can
  1015. 33:40learn anything." This is an explicit
  1016. 33:41call to update the skills. It used the
  1017. 33:44skill management tool. Remember in the
  1018. 33:46loop engineering, there's a tool called
  1019. 33:47skill management. And this is what it
  1020. 33:48did. And say, "Picape, I have
  1021. 33:50successfully created the video prep
  1022. 33:51skill. Let's find out about it." again.
  1023. 33:53Hermes skills media video prep. Double
  1024. 33:58click. It created this skill for me. It
  1025. 34:00says no m dashes spoken English define
  1026. 34:03jargon in line catchphrase. You can
  1027. 34:05build anything. You can learn anything.
  1028. 34:06Good. So we can test the loop a little
  1029. 34:08bit more. We can say use delegate task
  1030. 34:10to spawn two sub agents. One is
  1031. 34:12researching the LM eval harness. Another
  1032. 34:15one researching VLM architecture. The
  1033. 34:18agents was calling the tool caught
  1034. 34:19delegate task and successfully spawned
  1035. 34:21two background agents. So you can see
  1036. 34:23it's processing right now over here.
  1037. 34:25They're working in parallel.
  1038. 34:27Cool. Eventually return the results to
  1039. 34:30me. So you can see that it's spawn up
  1040. 34:31some sub agents to do task for me in
  1041. 34:33parallel. We've already tried the chrome
  1042. 34:35job. And the next one is spawn sub agent
  1043. 34:38that uses claw CLI. This is my favorite.
  1044. 34:40So this is my favorite. Spawn a sub
  1045. 34:42agent that uses the claw CLI in headless
  1046. 34:44mode. basically using claude to build a
  1047. 34:46Python script fetching the top five
  1048. 34:48hacker news stories to markdown and
  1049. 34:50verify it runs. Yep, it firstly
  1050. 34:52delegated a task to claude CLI is using
  1051. 34:56the skill. You can see remember we read
  1052. 34:58the skill here in autonomous AI agent
  1053. 35:00clot. It should be reading this skill.
  1054. 35:02Is this helpful? Let me know if this is
  1055. 35:03helpful. I mean, I was trying to show
  1056. 35:05you more examples in the walkthrough so
  1057. 35:07that it feels more concrete because a
  1058. 35:09lot of people were asking me for
  1059. 35:10implementations from my previous videos
  1060. 35:12and I feel like Hermas is really good
  1061. 35:14example.
  1062. 35:17Okay, it created this uh code for me and
  1063. 35:20used the cloud. It created the fetch
  1064. 35:23hack news.py and just say show me the
  1065. 35:27result after
  1066. 35:30running the script.
  1067. 35:34Yep. So these are the top five hacker
  1068. 35:36news stories on hacker news. Let's come
  1069. 35:38check it's correct. Remember what
  1070. 35:39happened was that we sent the question
  1071. 35:41from the gateway. It went through the
  1072. 35:42agent run and the LLM basically caught
  1073. 35:45up the tool called delegate task.
  1074. 35:47Delegate task delegated the task to sub
  1075. 35:49agent and the sub agent was running claw
  1076. 35:51code and cl code wrote that script and
  1077. 35:53sent it back and then this agent run the
  1078. 35:56script again and then we get the
  1079. 35:58results. This was the harness in the
  1080. 35:59loop engineering in this entire thing. I
  1081. 36:01don't think it self-t triggered
  1082. 36:03self-learning skills yet. It only
  1083. 36:06triggered learning the memory. I'm just
  1084. 36:08going to ask explicitly. Did you save
  1085. 36:09any skills by yourself or did you only
  1086. 36:11save memory so far? Oh yeah, but we
  1087. 36:13created that. But we created
  1088. 36:16that skill. You did not self realize
  1089. 36:21you need to save skill. You should
  1090. 36:25proactively offer to save a skill after
  1091. 36:27difficult iterative task with five or
  1092. 36:29more tool calls. But so far we haven't
  1093. 36:31done anything difficult. So we can test
  1094. 36:32the gateway. I can summarize what we
  1095. 36:35worked on today. H this chrome job
  1096. 36:37failed. It didn't tell me 10 jokes. It
  1097. 36:39only told me two. What were the rest of
  1098. 36:43the jokes? Don't lie to me because maybe
  1099. 36:46you just created it. Anyways, Hermes
  1100. 36:48team, if you're watching this, your
  1101. 36:49chrome job is not updating it unless I
  1102. 36:53ask for the results. So maybe this is
  1103. 36:55something you guys can fix. All right,
  1104. 36:56let's move on and summarize what we
  1105. 36:58worked on today because in this current
  1106. 37:00chat, it doesn't know the rest of the
  1107. 37:03stuff here. So, let's see if it's
  1108. 37:04calling the actual remember we say state
  1109. 37:07DB, memory MD, and skills.m MD. Good.
  1110. 37:10Actually knows. Okay, it says jokes on
  1111. 37:12demand system check on the OS version,
  1112. 37:15which is what we did. YouTube stats.
  1113. 37:17Yes, we did that. Hacker news. Okay,
  1114. 37:19good. WhatsApp here is just a gateway as
  1115. 37:22an interface. It's doing the same thing
  1116. 37:24as the Hermist agent desktop. This is
  1117. 37:26very cool. I mean, this is similar to
  1118. 37:28OpenC Claw, but still every time when I
  1119. 37:30feel I can just have a personal
  1120. 37:31assistant handy on my WhatsApp. This is
  1121. 37:33really cool. And imagine if I host this
  1122. 37:35on a virtual machine and it runs
  1123. 37:37forever. I can just basically ask
  1124. 37:39WhatsApp what's going on and it's going
  1125. 37:41to reply to me. Ooh, I completely
  1126. 37:42forgot. You should set up WhatsApp by
  1127. 37:44typing in her WhatsApp and then just
  1128. 37:46follow the instructions to set it up.
  1129. 37:48Should be straightforward. Last but not
  1130. 37:50least, where is Eval? All right, where
  1131. 37:52is our LM ops? Remember in our original
  1132. 37:55chart we had this blue box on the right
  1133. 37:56hand side called LM ops. So
  1134. 37:59unfortunately on Hermas based on my
  1135. 38:01current research I don't think it has an
  1136. 38:03LM ops or an Eval system. You probably
  1137. 38:06need to build it yourself but it does
  1138. 38:08trace the run. There's a thing called
  1139. 38:10trajectory export and logs and it's
  1140. 38:12basically just logging the whole thing
  1141. 38:13but there's no eval. If you watched our
  1142. 38:15previous video, the eval basically you
  1143. 38:16can use tools like land smith land view
  1144. 38:19fuse and all these things that can track
  1145. 38:21the entire agent run and some events
  1146. 38:23that happen like tool calls. How many
  1147. 38:25times did you call LLMs? Did you use
  1148. 38:27some cheaper models to summarize things
  1149. 38:29into your semantic memory?
  1150. 38:30Unfortunately, Hermas doesn't do that.
  1151. 38:32I'm curious why. Maybe because it's a
  1152. 38:34local system, so you can just customize
  1153. 38:36it yourself. It's under Hermas. You can
  1154. 38:38click into logs. You can see all the
  1155. 38:40logs here. Can see the errors. can see
  1156. 38:42the agent runs. Gateway logs to gateway
  1157. 38:45is the entry point which is WhatsApp or
  1158. 38:47this desktop. I hope this was helpful. I
  1159. 38:49think there's quite a huge amount of
  1160. 38:52work that we covered today which are
  1161. 38:54exactly what claw code can already do or
  1162. 38:56open claw can already do. I feel that
  1163. 38:58what makes her special at least compared
  1164. 39:00to claw code is that it's self updating
  1165. 39:03all these things locally. And maybe this
  1166. 39:05is good for privacy reasons because I
  1167. 39:07don't see why we should not save it on
  1168. 39:10the cloud. Other than that, everything
  1169. 39:11else should be very similar, right? It
  1170. 39:12runs the loop every time when the agent
  1171. 39:14needs to do some task uh kind of call
  1172. 39:15all these tools and you can even
  1173. 39:17schedule chrome drop and then the
  1174. 39:18memories and the skills are getting auto
  1175. 39:21updated uh into the procedure memory,
  1176. 39:23semantic memory and the state data will
  1177. 39:25be saved into the episodic memory. So I
  1178. 39:27would say this is a pretty standard
  1179. 39:28harness for an agent implementation.
  1180. 39:31Nothing too fancy, but I'm sure they did
  1181. 39:33a ton of work. But I feel like this is a
  1182. 39:35good example for how builder these days
  1183. 39:37should be building products because this
  1184. 39:39is becoming a standard, right? Every AI
  1185. 39:42agent tool should be self-improving and
  1186. 39:44self- evvolving. And depending on the
  1187. 39:46user request, you can either keep it
  1188. 39:48local, you can keep it on the cloud,
  1189. 39:49that's up to you. I feel like the skill
  1190. 39:51and memory accumulation is probably the
  1191. 39:54most valuable thing in this because the
  1192. 39:56user kind of just over time get locked
  1193. 39:58in and it remembers who I am. Which is
  1194. 40:01why like recently I don't switch to
  1195. 40:03other AI anymore. I just use cl code. It
  1196. 40:05has a semantic memory about who I am, my
  1197. 40:07company information, ultimately
  1198. 40:11YouTube videos. But yeah, I think it's a
  1199. 40:13cool framework. You guys should try it
  1200. 40:15out. It's really fun. Cool. I hope this
  1201. 40:17video was helpful. If you have any
  1202. 40:18questions, please leave down a comment.
  1203. 40:20And if you enjoyed it, please give me a
  1204. 40:22like and subscribe. Let me know what
  1205. 40:23else you would like to watch. I will see
  1206. 40:25you next time. Thanks. Hey everyone,
  1207. 40:26this is Sean. So today, let's talk about
  1208. 40:29loop versus graph engineering. Two
  1209. 40:31biggest buzzwords in AI agent space
  1210. 40:33recently. I am probably exactly like
  1211. 40:36most of you guys, which is I don't know
  1212. 40:38how to keep up with all these new
  1213. 40:39concepts, but I still find it very
  1214. 40:41interesting because whenever a new
  1215. 40:43concept becomes viral, usually I'm
  1216. 40:45curious about what triggered it. What I
  1217. 40:47find out is that this guy called Peter
  1218. 40:48Steinberger who's also the author of
  1219. 40:50this little lobster which is open claw
  1220. 40:53said that are we still talking about
  1221. 40:55loops or did we shift to graphs yet? And
  1222. 40:58he posted this on July 18th 2026 which
  1223. 41:01is about 13 days ago and he got 3
  1224. 41:04million views and look at this first
  1225. 41:06comment. It said bro stop I am on
  1226. 41:08vacation. So sometimes you know these
  1227. 41:11important people in the space who
  1228. 41:13mention a keyword and then everybody's
  1229. 41:16going to start talking about it. I
  1230. 41:17remember exactly something like that in
  1231. 41:19context engineering. And for today this
  1232. 41:21video we are going to demystify all of
  1233. 41:23these core AI agent concepts and
  1234. 41:25especially on loop and the graph. We're
  1235. 41:28going to jump in a little bit on the
  1236. 41:29system design real quick. And at the
  1237. 41:31same time, we're also going to walk you
  1238. 41:32through a real coding example uh called
  1239. 41:35Waku- Agent, which is under my GitHub,
  1240. 41:37Shen Sean Chen, and we published this
  1241. 41:39about two three weeks ago, and we got
  1242. 41:41more than 700 stars. If you think this
  1243. 41:43video is helpful for you, please give us
  1244. 41:44a star, give us a like, and that would
  1245. 41:46be very helpful for us. With Waku As
  1246. 41:48agent, you can launch a dashboard like
  1247. 41:49this, and you can just add something
  1248. 41:50like what's up on my Google calendar on
  1249. 41:56Thursday,
  1250. 41:58right? And then it's going to check what
  1251. 42:00kind of graph it's going to use, decides
  1252. 42:02if it needs to retrieve some memories,
  1253. 42:04and then eventually decide, you know,
  1254. 42:06what kind of tools it should be using.
  1255. 42:07Right? Right here, it's listing the
  1256. 42:08events from my Google calendar and
  1257. 42:10eventually send you the reply, and it
  1258. 42:13did a bit of a testing in the vow
  1259. 42:15system. And then, you know, it spit up
  1260. 42:17the gave me the final answer for the
  1261. 42:20results. I'm going to talk a bit more
  1262. 42:21details into this, but without further
  1263. 42:23ado, let's get started on the system
  1264. 42:24design. First thing first, I want to
  1265. 42:26talk about the AI agent engineering
  1266. 42:29ladder that we have come across over the
  1267. 42:31past three years because I think that
  1268. 42:33helps us to kind of understand where we
  1269. 42:36came from and where we're going. And the
  1270. 42:38first one is obviously prompt
  1271. 42:40engineering. And to me, prompt
  1272. 42:41engineering was the very first or early
  1273. 42:44version of how people start to get used
  1274. 42:46to how to use an LLM. I feel the most
  1275. 42:49accurate word to describe this was role
  1276. 42:51playing because at the very beginning,
  1277. 42:52we didn't know how to use LMS. We had
  1278. 42:54chatbt 3.5 and we tell it, hey, you are
  1279. 42:58the best poet like your Shakespeare or
  1280. 43:01Leai, right? Write me a poem about this
  1281. 43:04scenery I'm looking at right now in
  1282. 43:06Switzerland.
  1283. 43:07Then it's going to pretend that it's one
  1284. 43:09of those poets you just mentioned and
  1285. 43:11then talk to you back. Right? That was
  1286. 43:12prompt engineering. You basically draft
  1287. 43:14the prompt so that you control the LM to
  1288. 43:17talk to you in a certain way you want.
  1289. 43:19And then we quickly evolved into what we
  1290. 43:22call context engineering. And that is
  1291. 43:24because when people move from just
  1292. 43:28playing with LM as a consumer product to
  1293. 43:31actual workflows, you realize that you
  1294. 43:33not only need prompt but also you need
  1295. 43:35to feed in data. For example, if you
  1296. 43:38build a customer service chatbot or if
  1297. 43:40you build a sales agent like Automanas,
  1298. 43:43which is my company, you need to really
  1299. 43:44think about how do you construct the
  1300. 43:46context for the agent so that the agent
  1301. 43:49will be able to talk to clients on
  1302. 43:51behalf of the businesses, right? So, you
  1303. 43:54not only will say, hey, you are the best
  1304. 43:57salesperson in the world. You will talk
  1305. 43:59in certain ways. Do not say certain
  1306. 44:01things. You also need to feed in data
  1307. 44:03such as what are some of the what are
  1308. 44:04some of the customer relationship data
  1309. 44:06people already have on their Excel
  1310. 44:08sheets, Google sheets, CRM system, all
  1311. 44:11these kind of things and you need to
  1312. 44:12really put them together and make sure
  1313. 44:14that the agent can accurately be fed
  1314. 44:17with the right information. So we moved
  1315. 44:19from prompt to context very quickly to
  1316. 44:21fix that problem. And then we moved on
  1317. 44:23to skills which I think is sort of
  1318. 44:26teaching the AI what kind of procedure
  1319. 44:29it should follow. Right? It's a
  1320. 44:31procedure following process in AI agent
  1321. 44:34hardness. This is called procedural
  1322. 44:36memory. It's a memory that you tell say
  1323. 44:38you tell a kid, hey, when you walk home,
  1324. 44:42walk on the right hand side if you're in
  1325. 44:44the US or in China, right? Walk on the
  1326. 44:46left hand side if you're in the UK or
  1327. 44:48Japan. That is a procedural memory that
  1328. 44:51you want a person or you want an LLM to
  1329. 44:54remember. And there's no, you know,
  1330. 44:57extra data. It's just fact that it
  1331. 44:59should be doing when a certain situation
  1332. 45:01happens. Okay. So why do we need skills?
  1333. 45:04It's because that if you just p if you
  1334. 45:07if you provide a lot of context to the
  1335. 45:09AI, sometimes it can be a little
  1336. 45:11repetitive. You don't want to always
  1337. 45:14provide you know the same order to AI
  1338. 45:16again and again and again. So having a
  1339. 45:18skill to determine a workflow becomes
  1340. 45:21really really handy. Just like for
  1341. 45:24example if you're coding on claw code
  1342. 45:26you don't want to explain to claude that
  1343. 45:29do not use emoji do not use emoji do not
  1344. 45:31use emoji or I prefer to use emoji use
  1345. 45:34these emojis don't use the brain emoji I
  1346. 45:36hate that right these kind of things you
  1347. 45:38have to repeatedly tell the context then
  1348. 45:40it's becoming less convenient so you can
  1349. 45:42build up skill to make sure the LM is
  1350. 45:45following that procedure then comes with
  1351. 45:47loop what does loop do a loop is
  1352. 45:50basically saying hey Maybe sometimes you
  1353. 45:53have a goal and you want to finish that
  1354. 45:56goal but we don't know the exact skills
  1355. 46:00you should have to finish that goal. We
  1356. 46:02might tell you, hey, there are a bunch
  1357. 46:03of tools you can use, right? That's why
  1358. 46:05Enthropic come up with MCPS, you can
  1359. 46:07call your Google calendar, you can call
  1360. 46:09your Gmails APIs, you can call your
  1361. 46:11GitHub APIs. These tools become handy
  1362. 46:14and then you're telling the LM be like,
  1363. 46:16okay, run a loop, know your goal, which
  1364. 46:19is maybe help me fix this bug that the
  1365. 46:21customer come up with. Run this loop and
  1366. 46:24here are a bunch of tools. Here's a web
  1367. 46:26search tool. Here's a bunch of API
  1368. 46:27calls. Here's a bunch of MCPS. Use them.
  1369. 46:31loop it until at some point you finish
  1370. 46:33the goal and that's the end of the
  1371. 46:35iteration. And then the question comes
  1372. 46:38why do we need graph? What does graph
  1373. 46:41do? Because graph technically is a
  1374. 46:43procedure that is predetermined. Many
  1375. 46:46people are criticizing graph engineering
  1376. 46:47is not new because maybe in 2023 I
  1377. 46:50remember people were already using
  1378. 46:51airflow. people using step functions.
  1379. 46:53People are talking about how do we make
  1380. 46:55sure that a deterministic workflow can
  1381. 46:58be set up properly so that we don't just
  1382. 47:00tell LLM to make all the decisions
  1383. 47:01because sometimes we know exactly how a
  1384. 47:03certain task needs to be done. There's a
  1385. 47:04step there. There's an SOP there. So to
  1386. 47:08me it kind of feels like okay we moved
  1387. 47:10from skills which is strict procedure
  1388. 47:13following to loops which is hey agents
  1389. 47:15go figure it out yourself to eventually
  1390. 47:18we realize that hey we need a mixture of
  1391. 47:20both. That in my opinion is graph. Okay.
  1392. 47:24Technically sometimes you should write
  1393. 47:26the skill first. Okay. How do you cross
  1394. 47:30the road? On what side do you walk in
  1395. 47:32the pavement depending on which country
  1396. 47:33you're in? And do you respond to a
  1397. 47:36client? How do you respond to me when
  1398. 47:38I'm coding with an coding agent? Right?
  1399. 47:40And then you should turn that into a
  1400. 47:41graph. If you see that there's a lot of
  1401. 47:43repetition in the workflows, especially
  1402. 47:45if you're building a workflow for
  1403. 47:46corporate, right? If you're working in
  1404. 47:48uh e-commerce, maybe you see exactly how
  1405. 47:51your customer service should be
  1406. 47:53answering questions related to
  1407. 47:54logistics, to refund, to checking
  1408. 47:56samples, you realize that you can really
  1409. 47:59consolidate this into a graph. When the
  1410. 48:01workflow stops changing, perhaps there's
  1411. 48:03still part of the workflow that you need
  1412. 48:04to use a loop where the loop is
  1413. 48:07basically doing this explorative work
  1414. 48:09out there, right? It's probably doing,
  1415. 48:11you know, a bunch of research for you.
  1416. 48:13Um I think deep research is one of the
  1417. 48:15early examples of a loop engineering
  1418. 48:17workflow here where you just say hey go
  1419. 48:19crazy just go search the internet I want
  1420. 48:21to report these kind of work are you
  1421. 48:24know less standardized or it doesn't
  1422. 48:26have an SOP in it. It's just about I
  1423. 48:29want more information or I have a
  1424. 48:31certain goal use the tools available to
  1425. 48:34you to figure it out for me. is quite
  1426. 48:36different from graph because sometimes
  1427. 48:38we know exactly what tools they should
  1428. 48:40be using to uh finalize a task for me.
  1429. 48:43Okay guys, that was a conceptual
  1430. 48:46walkthrough of this AI agent engineering
  1431. 48:48ladder. Now let's take a look at what a
  1432. 48:50loop and a graph look like. What a loop
  1433. 48:53does is that it discovers what to do
  1434. 48:55next. So maybe this is you and you ask a
  1435. 48:58question to an LLM and the LM is
  1436. 49:00basically saying, "Hey, do I need to use
  1437. 49:02any tools? If yes, choose a tool, run
  1438. 49:05it, and check if it finish the task,
  1439. 49:08come back to the LM, and then loop it
  1440. 49:10again and again and again until at some
  1441. 49:12point LM is like, hey, we don't need the
  1442. 49:14tool anymore. Let's reply to the user.
  1443. 49:17Examples here could be like you are
  1444. 49:19trying to fix a bug on a pull request on
  1445. 49:22GitHub. You basically ask claw code or
  1446. 49:24codeex and be like, fix this bug, tell
  1447. 49:27me when it's fixed. And it's going to go
  1448. 49:28ahead and then use a bunch of tools like
  1449. 49:31web search, github CLI, checking your
  1450. 49:33superbase, checking your AWS, checking
  1451. 49:35your Google cloud, all these kind of
  1452. 49:36stuff and then at the end saying okay
  1453. 49:38we're done. Okay, bug is fixed because
  1454. 49:41of ABCD that is a loop. A graph on the
  1455. 49:44other hand
  1456. 49:46is somewhat similar but uh it's not
  1457. 49:48exactly the same. So you might have a so
  1458. 49:50you might have a standardized process
  1459. 49:52every day and be like oh I want to
  1460. 49:54understand how many people submitted
  1461. 49:55pull requests overnight and who are
  1462. 49:57these guys who submitted some tasks can
  1463. 50:00we take a look at them and uh tell me if
  1464. 50:03uh things are fixed or not and maybe at
  1465. 50:06the same time I want to understand you
  1466. 50:08know what are some of the latest news
  1467. 50:10out there on AI agents and maybe we want
  1468. 50:12to do some web search as well maybe we
  1469. 50:14want to run some you know git commands
  1470. 50:16at the same time we know exactly how it
  1471. 50:18works Okay. And then you probably want
  1472. 50:21this to be done in parallel. So you ask
  1473. 50:24a question and then it's going to check
  1474. 50:25out the GitHub to check out the pull
  1475. 50:27request. It's going to search the
  1476. 50:28website for you to do some explorative
  1477. 50:31analysis. And it's probably also going
  1478. 50:33to check your calendar plus the memories
  1479. 50:35stored locally or on the cloud. And then
  1480. 50:39it's going to synthesize these
  1481. 50:40information and tell us, hey, is there
  1482. 50:42anything else to do? If not, end this
  1483. 50:44and reply. If yes, please explain what
  1484. 50:48kind of things do you still need to do.
  1485. 50:49Okay, you see the difference here. So
  1486. 50:51sometimes you can have some loops here.
  1487. 50:53Okay, maybe having an agent maybe having
  1488. 50:55an agent loop to search the web or have
  1489. 50:58an agent loop to fix the bugs on GitHub
  1490. 51:00is part of this graph. Okay, so you can
  1491. 51:03see graph is basically saying, okay, I
  1492. 51:05know exactly what you should be
  1493. 51:06checking. Maybe they're in parallel,
  1494. 51:08maybe they are happening in sequences.
  1495. 51:10Do it the way I want. And uh once you
  1496. 51:13finish uh synthesize it and tell me the
  1497. 51:15answer. Are you guys still with me?
  1498. 51:17Let's check a real example first. Come
  1499. 51:19back here. Let's come to Waku agent
  1500. 51:22dashboard. The way you set it up, by the
  1501. 51:23way, is come to this website
  1502. 51:25github.com/jennonchan/wacu-agent.
  1503. 51:28You can either click on code and then
  1504. 51:30copy this and then type into your
  1505. 51:33terminal and say get clone and paste
  1506. 51:36that in and hit enter which is this way.
  1507. 51:38And you can just copy this and paste in
  1508. 51:39your terminal. or recently we released a
  1509. 51:42new package in Python called Waku-
  1510. 51:44Agent. All you need to do is copy this,
  1511. 51:47pip install the Wacu agent into your
  1512. 51:48terminal, set up your environmental
  1513. 51:50keys, and then you can launch a
  1514. 51:51dashboard. Copy this,
  1515. 51:54paste this in
  1516. 51:58because I'm already using port 777. So,
  1517. 52:00let's use 778 as an example. Paste that
  1518. 52:04in. You can see this is it. Okay, since
  1519. 52:07I already set it up, I'll come back to
  1520. 52:08this localhost 7777.
  1521. 52:11So, what you're seeing right here is is
  1522. 52:12an entire AI agent harness starting from
  1523. 52:15the gateway, which can be the chat of
  1524. 52:17here, or you can use some other channels
  1525. 52:19such as Discord, Telegram, WhatsApp,
  1526. 52:21stuff like that. And then it's going to
  1527. 52:24check out a retrieval gate to see if we
  1528. 52:26need any procedural memory which is
  1529. 52:28skills as we mentioned or semantic
  1530. 52:30memory or episodic memory which are
  1531. 52:32durable facts or the dated events that
  1532. 52:34happened in your local memories. Okay.
  1533. 52:37And then it's going to run through an
  1534. 52:38agent loop using LM agents and calling
  1535. 52:40the tools and eventually give you the
  1536. 52:42reply during which the LM ops is going
  1537. 52:44to trace the data test it and then
  1538. 52:47release it the new version of the prompt
  1539. 52:49and feed back to the harness. This is
  1540. 52:51the entire harness. Okay. And loop
  1541. 52:53engineering is happening here in this
  1542. 52:54little loop. A graph workflow is
  1543. 52:57basically after this gateway you can
  1544. 53:00predefine some of the process in between
  1545. 53:02here. So we have a new tab called graph.
  1546. 53:04We currently have two graphs here. One
  1547. 53:05is called triage another one is called
  1548. 53:07gather. For triage graph what it does is
  1549. 53:10that it's saying okay start point and
  1550. 53:12then it's going to classify if it
  1551. 53:15requires some agent calls and then at
  1552. 53:16the same time it might just check out my
  1553. 53:18calendar and see what's going on. Okay.
  1554. 53:20And I can just be like what's up today?
  1555. 53:25All right, you can see it triggered this
  1556. 53:27triage first, right? It checked the
  1557. 53:29Google calendar for me and also at the
  1558. 53:31same time it was checking, you know, if
  1559. 53:32we need to have some uh need to do some
  1560. 53:35serious Asian calls here and it
  1561. 53:38eventually decided that okay, it's going
  1562. 53:39to use some tools. So they use the list
  1563. 53:41events read Apple calendar for me. So
  1564. 53:43you just saw a very simple graph kind of
  1565. 53:45call already. Okay. So what it did was
  1566. 53:48that it checked as I mentioned if it
  1567. 53:51needs some serious agent calls and in
  1568. 53:54parallel at the same time it was already
  1569. 53:55checking my calendar because I was
  1570. 53:56asking what's up today and then after
  1571. 53:58that use this loop to run to call the
  1572. 54:00tools like list events read apple
  1573. 54:02calendar and all these kind of stuff and
  1574. 54:04then uh give me the reply. Okay but what
  1575. 54:07if I need to test something else. So
  1576. 54:09here in this local Asian harness what I
  1577. 54:11recently updated is that you can very
  1578. 54:13clearly mention the workflow you built
  1579. 54:16in the graph we have gathered. So you
  1580. 54:18can clearly say slashgather and then say
  1581. 54:22tell me what's up with waku agents and
  1582. 54:24any new PRs any competitive projects.
  1583. 54:27Let's see what happens. You see that it
  1584. 54:29triggered this graph instead. Okay let's
  1585. 54:32come back to the overview. It's showing
  1586. 54:33me that it was trick triggering this
  1587. 54:35graph and it used these tools
  1588. 54:38simultaneously
  1589. 54:40and then it's synthesizing the answers
  1590. 54:42for me. Okay. So, so what it did was
  1591. 54:45that it gave me a morning brief on what
  1592. 54:48kind of PRs are out there and we have
  1593. 54:51released a new package which you can
  1594. 54:53check in this GitHub if you scroll down
  1595. 54:55a little bit. We have released a new
  1596. 54:56package for Asian graphs recently and
  1597. 54:59also there are some PRs to be reviewed.
  1598. 55:01Okay, let's come here. So, if you click
  1599. 55:04into pull request, you can see there are
  1600. 55:06some PRs here for me to be reviewed. And
  1601. 55:08for the web search, you can see that it
  1602. 55:10researched about Harrison Chase publish
  1603. 55:12your harness your memory and your own
  1604. 55:14videos in harness eval ranking. Okay, it
  1605. 55:18was doing the command research for me as
  1606. 55:20well. All right, so you might be
  1607. 55:22wondering how did this work? If we come
  1608. 55:25back to the tab for graph, you can see
  1609. 55:27that we clearly defined two different
  1610. 55:28use cases of graph. One is triage,
  1611. 55:30another one is gather. And they have
  1612. 55:32very specific ways of running the
  1613. 55:35agents. And they both use loops because
  1614. 55:38when it does the web search, it needs to
  1615. 55:39search for the information and come back
  1616. 55:41to me. Right? When it does the check
  1617. 55:43calendar, it's going to check the
  1618. 55:44calendar until it find out what's going
  1619. 55:46on in my calendar and then come back to
  1620. 55:48me. Okay? These things are all kind of
  1621. 55:50intertwined in this agent harness
  1622. 55:52system. So if we come back to the system
  1623. 55:53design, there's something interesting I
  1624. 55:55wanted to show you which is how exactly
  1625. 55:57did the workflow finish from the
  1626. 55:59beginning to the end from a time usage
  1627. 56:02and parallel processing perspective. For
  1628. 56:04the loop, an agent loop, for example,
  1629. 56:06web search is going to decide what it's
  1630. 56:08going to do first, right? It's going to
  1631. 56:10check out my GitHub calling tools one by
  1632. 56:13one, right? Maybe after the first call
  1633. 56:15the agent decided okay we need to do one
  1634. 56:17more call and then it called it again
  1635. 56:19and then it checked the calendar and
  1636. 56:21then it synthesized information for me.
  1637. 56:23But in the graph example which is this
  1638. 56:26which is this workflow for gathering
  1639. 56:29information. It used multiple tools in
  1640. 56:31parallel because we already told it you
  1641. 56:32need these four tools, right? Maybe each
  1642. 56:34one of them spend different amount of
  1643. 56:36time and it did the GitHub first and the
  1644. 56:38web search did the most amount of time.
  1645. 56:40Calendar check and memory retrieval did
  1646. 56:42the least amount of time because it's an
  1647. 56:44instant or maybe because it did not need
  1648. 56:46any memory retrieval at all. And then
  1649. 56:48it's going to synthesize the information
  1650. 56:50because it's got a lot of context at the
  1651. 56:52same time. So you cannot say so you
  1652. 56:54cannot say graph engineering is
  1653. 56:56replacing loop engineering because
  1654. 56:58they're coexisting and you cannot say
  1655. 57:00loop engineer graph engine is better
  1656. 57:01than the other one because we need them
  1657. 57:03in different use cases. A loop is
  1658. 57:05something you need when the model
  1659. 57:06decides what to call one step at a time.
  1660. 57:09A graph is something like when you know
  1661. 57:11the shape and you just want them to move
  1662. 57:14together. Okay, it really depends on the
  1663. 57:16situations you're building for. In your
  1664. 57:18local files, we have a WKU folder and
  1665. 57:21then we have a graph folder and within
  1666. 57:24the graph folder, we have defined the
  1667. 57:27graph engines which is literally a graph
  1668. 57:31and you're going to add nodes which in
  1669. 57:33our case could be web search tools,
  1670. 57:35could be agent calls, could be MCPS,
  1671. 57:37anything right and then you're also
  1672. 57:39defining edges which is okay do we where
  1673. 57:41do we go after using one tool right
  1674. 57:44after the web search do we go
  1675. 57:45synthesizing or do we go somewhere else
  1676. 57:47Right graph we also defined nodes which
  1677. 57:51is basically saying what kind of things
  1678. 57:53can be a node right to calls LLM calls
  1679. 57:57agent calls router these kind of things
  1680. 58:00and then under this graph folder we also
  1681. 58:02have another folder called workflows you
  1682. 58:04can see we have triage here right
  1683. 58:06remember triage was the chart we showed
  1684. 58:08you which is it's going to classify if
  1685. 58:10it needs to use some complex models to
  1686. 58:11do agent calls or is it doing some
  1687. 58:14calendar checking at the same time these
  1688. 58:16two things are happening in parallel at
  1689. 58:18the same time. And the close is actually
  1690. 58:20very simple. You define these
  1691. 58:22functionalities to make sure that the
  1692. 58:24triage graph is working properly with
  1693. 58:26this predefined workflows, right? You
  1694. 58:29add node when you need to add more
  1695. 58:31tools. You add edges by defining the
  1696. 58:34initial problem from start can go to
  1697. 58:36either classify or check calendar at the
  1698. 58:39same time. If we check out the gather
  1699. 58:42tool as well, remember it's a similar
  1700. 58:44thing as our preview chart. You can scan
  1701. 58:46the GitHub website calendar memory.
  1702. 58:50Okay, GitHub website calendar memory.
  1703. 58:53These are all the nodes and after that
  1704. 58:55you synthesize it and then you decide if
  1705. 58:57we should return the answer. Obviously
  1706. 58:59we're doing the same thing. We're
  1707. 59:00building the graph here, right? We're
  1708. 59:02adding these nodes and we are adding
  1709. 59:05those edges too. I hope this is easy to
  1710. 59:07understand and you can feel free to add
  1711. 59:11more.py files here to make this graph
  1712. 59:13workflows even larger. And in that case,
  1713. 59:16I would love to work with you and then
  1714. 59:17you can feel free to contribute to this
  1715. 59:19repo and be one of our contributors and
  1716. 59:22submit your own workflows because once
  1717. 59:24you submit your own workflows and it's
  1718. 59:26approved, it will be available here. If
  1719. 59:28I type in slashgraphs,
  1720. 59:31it's going to tell me that it has
  1721. 59:34gather, it has triage and triage is the
  1722. 59:36router. It's basically going to read the
  1723. 59:38local files of these workflows from
  1724. 59:40graphs and people can use it. This will
  1725. 59:43be very cool and then we'll probably
  1726. 59:44become a community if you guys are
  1727. 59:46interested and feel free to be our
  1728. 59:47contributors and also if you're
  1729. 59:49interested in talking more in depth of
  1730. 59:51these concepts with us if you feel you
  1731. 59:53want to discuss with me about your
  1732. 59:54workflows and any questions you have
  1733. 59:56about system design about these agent
  1734. 59:57hardness and all these kind of stuff you
  1735. 59:59can feel free to come to my personal
  1736. 1:00:00website shantchan.io O and then click on
  1737. 1:00:03join over here to join our community.
  1738. 1:00:05I'm just getting started to do this
  1739. 1:00:07because I don't have enough time to
  1740. 1:00:09answer all of your questions. Feel like
  1741. 1:00:10the most efficient way is that I can
  1742. 1:00:12host some live sessions with you uh
  1743. 1:00:13twice a month so that I'll be able to
  1744. 1:00:15answer most of your questions and we can
  1745. 1:00:17prepare some you know build sessions
  1746. 1:00:19real time showing you my real setup and
  1747. 1:00:21all this kind of stuff. If you join this
  1748. 1:00:22community, you will be assigned to a
  1749. 1:00:24private discord and I will share more
  1750. 1:00:27details there including the original
  1751. 1:00:29files of all of these system design
  1752. 1:00:31charts that I built in the past uh in
  1753. 1:00:34real code so that you can sort of open
  1754. 1:00:36it in your own scala draw website to
  1755. 1:00:38learn about it. So last but not least, I
  1756. 1:00:40think this is an important question we
  1757. 1:00:41should ask ourselves which is isn't this
  1758. 1:00:44just a deterministic workflow from 2023?
  1759. 1:00:46What is new now is that some of the
  1760. 1:00:48nodes that we're using right now are not
  1761. 1:00:49deterministic. could be a LM call and at
  1762. 1:00:52the same time sometimes the model can
  1763. 1:00:54pick the edge right the routing is the
  1764. 1:00:56way that you know you're letting the LM
  1765. 1:00:59as a judge do I pick a simple model to
  1766. 1:01:02answer the questions or do I use a more
  1767. 1:01:05complex model to answer this question
  1768. 1:01:07and also you need some guards which a
  1769. 1:01:09previous directional graph scheduler
  1770. 1:01:12never did these are buzzwords as I
  1771. 1:01:14mentioned but buzzwords are viral for a
  1772. 1:01:16reason sometimes because of famous
  1773. 1:01:18people sometimes because it's actually
  1774. 1:01:19useful So I hope this kind of videos is
  1775. 1:01:22helpful for you and again if you have
  1776. 1:01:24any questions feel free to ask me and I
  1777. 1:01:27would love to answer your questions live
  1778. 1:01:28in our community just come to my
  1779. 1:01:30personal website shanten.io and then
  1780. 1:01:32there are a bunch of sources here
  1781. 1:01:34looking forward to and if you love this
  1782. 1:01:36project please give us a star on GitHub
  1783. 1:01:37repo and I would love to work with you.
  1784. 1:01:40Thank you so much for your attention.
  1785. 1:01:42Appreciate it.

About this transcript

This page contains the full transcript of 8 сентября 2026 г. by Andrei, generated from the public captions YouTube serves with the video. The transcript has 12,711 words across 1,785 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.