8 сентября 2026 г. — Transcript
Full transcript
- 0:00So today we're going to go through
- 0:01probably the biggest buzzwords in AI
- 0:03agent system recently agent harness loop
- 0:06engineering ln ops which stands for
- 0:08large language models operations eval
- 0:11which stands for evaluation system for
- 0:13AI agents and these things become
- 0:15popular or become viral on the internet
- 0:17not because they are just some really
- 0:19complicated concepts instead they're
- 0:21actually very simple and I believe that
- 0:22simple building blocks will actually
- 0:24help us build the biggest architecture
- 0:26in the world that will function like an
- 0:28intelligent system let's walk through
- 0:29step by step and does not matter if
- 0:31you're technical or not. We're going to
- 0:33go through this and we'll make sure that
- 0:35you're equipped with the right knowledge
- 0:37for prompting your way through building
- 0:39such a system in the future. Let's jump
- 0:41in and get started. For those of you who
- 0:42have watched my previous video on AI
- 0:44agent memories, you're probably already
- 0:46familiar with this chart. This is an AI
- 0:48agent run, which means that it takes an
- 0:50input from a user prompt. For example,
- 0:52you're asking Chad GPT or DeepSeek a
- 0:55question and say, "Hey, when was Sam
- 0:56Alman fired from OpenAI?" And then he's
- 0:58going to go through entire run. But the
- 1:00end goal is that you want to get a
- 1:01response. This is actually ephemeral
- 1:03which means that there's no memory in
- 1:04this at all. We're sending that question
- 1:06when was Saman fired and any chat
- 1:08history that's currently in the chat.
- 1:10For example, maybe we had some
- 1:11conversations before that which for
- 1:13example could be you should talk to me
- 1:15like Elon Musk grilling on Samman
- 1:17because they don't like each other. And
- 1:19then these things will be fed into this
- 1:20thing called a working memory or a
- 1:22context RAM. In this video, we probably
- 1:24won't dive too much in depth into the
- 1:26memory system because there's a previous
- 1:27video talking about it already. But I'll
- 1:29just quickly go through it and then
- 1:30we'll introduce the concept of what a
- 1:32harness means. When you have this kind
- 1:33of short-term working memory, there will
- 1:35be an LLM or a large language model
- 1:37which performs as a question and answer
- 1:39agent and at the end you're going to get
- 1:41a reply. But the problem with a simple
- 1:43agent run with simply just the question,
- 1:46current chat history and system prompt
- 1:48is that the memory is very shortterm.
- 1:50But when you run an AI agent system,
- 1:52sometimes we need extra memories. For
- 1:54example, how should the agent respond to
- 1:56the person? A procedural memory is
- 1:58exactly that. It basically tells the
- 2:00agent how to act and what are some of
- 2:01the instructions for this skill. We
- 2:03might also want the agent to know some
- 2:04durable facts about this context. For
- 2:07example, I might want to compare my own
- 2:09early stage startup journey with Sam
- 2:11Alman's early startup journey. We need
- 2:13this agent to have a memory of who I am,
- 2:15which in this context would be a durable
- 2:17facts or a semantic memory. Who Shawn
- 2:19is, what did he build in the past? These
- 2:21kind of things became a fact that you
- 2:23want your agent to know, but they're not
- 2:24publicly available if you're not famous
- 2:26because the AI model won't be trained on
- 2:28such information yet. But if you're
- 2:29famous already, you can skip this. They
- 2:31already know who you are. And another
- 2:32thing we need is called episodic memory.
- 2:34And they include things like the past
- 2:36events or past chat history that does
- 2:38not exist in this current conversation.
- 2:40For example, I might suddenly be
- 2:41wondering when was the last time I was
- 2:43preparing for a job application and can
- 2:45we retrieve that information and match,
- 2:47you know, if we can get a job in CHIGBT.
- 2:49So these things will be retrieved from
- 2:51this thing called an episodic memory,
- 2:53which is basically a time series of the
- 2:54previous conversations or previous
- 2:56triggers that happen if you have a more
- 2:58complex system. So for those of you who
- 3:00have watched my previous memory agent
- 3:02system design, you might be wondering,
- 3:03Sean, why are you repeating all of these
- 3:05things? And that is because if you think
- 3:07about the entire thing that we just
- 3:08covered in the past few minutes, we're
- 3:10really stating the one fact that a large
- 3:12language model can't do these things by
- 3:15itself. It's like a really powerful
- 3:16brain that knows everything about
- 3:19humanity, everything about science,
- 3:21anything that happened in human or
- 3:22biology history, but it does not know
- 3:25you. With you or the software who's
- 3:27running this AI agent system, the large
- 3:30language model has no clue with how you
- 3:32want it to perform. This is why the
- 3:35concept called harness becomes really
- 3:37important in this. What harness means
- 3:39literally is that it's a set of harness
- 3:41tools that you use to control a horse
- 3:44when you're doing a horse riding.
- 3:45Imagine this large language model is a
- 3:47horse, right? This horse is very
- 3:48powerful. They can run around, but if
- 3:50you don't have a good set of tools to
- 3:52ride this horse, you could just get
- 3:54hurt. You might go anywhere. You might
- 3:56go somewhere random. If you're in a war,
- 3:58you don't want that to happen. And
- 3:59that's why we're doing all of these to
- 4:01make sure we have good control over this
- 4:03large language model and make sure we're
- 4:05utilizing it at its maximum potential.
- 4:07That's why in addition to just the
- 4:09question or use a prompt and getting the
- 4:11reply and we're feeding them all in as a
- 4:13working memory which can be enhanced by
- 4:15these three memories we just talked
- 4:17about. And in order for these three
- 4:18memories to actually work, there's a bit
- 4:20more details and they're all included in
- 4:22harness. Remember hardness means we're
- 4:24building this Asian framework to control
- 4:27this large language model so that it
- 4:29works the way we want. For those of you
- 4:30who study statistics or machine
- 4:32learning, you would understand that a
- 4:33large language model is actually
- 4:35predicting the probability of the next
- 4:37word that it should spit out. When
- 4:38everything comes with probability,
- 4:40there's randomness in it. But when we
- 4:42solve problems, we sometimes don't want
- 4:44too much randomness. So that's why we
- 4:46need to have a good control over this
- 4:47technology. Now let's continue to finish
- 4:49this harness. There are lots of tools on
- 4:51the market that's already quite useful.
- 4:52For example, you could try uh tools like
- 4:54langraph, lane chain or pyantic and
- 4:56there are many others. In this video, we
- 4:58won't dive too much in depth into that
- 4:59and we're going to finish building up
- 5:01this harness before we move on to the
- 5:02next topic. So again, for this agent to
- 5:04work properly, we need this memory
- 5:05system to work. But this memory system
- 5:08needs an update system because memory
- 5:10doesn't just exist or pop up from
- 5:11nowhere. You need to constantly update
- 5:13it. That's why we need a database to
- 5:15store all these memories so that when
- 5:17the agent is running in this agent run,
- 5:19it knows where to retrieve these
- 5:21memories. So whenever you see an icon
- 5:23like this, this is a database. Okay,
- 5:25procedural memory is basically remember
- 5:27it's it's about instructions, right?
- 5:28It's about how the agent should be
- 5:30acting. It's like with a hardness on the
- 5:31horse, you want the horse to write
- 5:33faster or slower. Normally, these are
- 5:34just files or text. And that's why you
- 5:36probably heard of this buzz word called
- 5:37skills. skill is basically a piece of
- 5:40text in a markdown file that you feed
- 5:41into AI agent like claw code. But if you
- 5:44want to harness the system, well, just
- 5:46having files and text is not enough and
- 5:49they're stored in databases say like
- 5:51AWS, Superbase, Google Cloud, you know,
- 5:53Azure, all these kind of places or you
- 5:55can set up your own server at home if
- 5:56you want, but that's just too expensive.
- 5:58You don't want to do that. And in order
- 5:59for this harness to work properly, you
- 6:01also need to figure out how to store the
- 6:03memories. Okay? So for example, the
- 6:05episodic memory is the time series of
- 6:07the events that happened or the previous
- 6:08chat history. Again, the way we store it
- 6:10is actually very simple. You just track
- 6:12every single thing that happened and
- 6:14then it's going to become like a very
- 6:15long list of things that happen in
- 6:17history with timestamps. Durable fact is
- 6:19a different story. You can either input
- 6:20it yourself or you want the system to
- 6:22sort of automatically evolve over time.
- 6:24And the way for it to evolve is that you
- 6:26want to consolidate some of the
- 6:27conversations into the semantic memory.
- 6:30If I'm running a T2C brand e-commerce
- 6:32company, perhaps my customers have
- 6:33talked to my customer service agents for
- 6:35a million times about how do I get
- 6:37reimbured if this product does not work.
- 6:39You want to consolidate these
- 6:40conversations and distill them as a fact
- 6:43into the semantic memory. And that's why
- 6:45here we have a little gate here. If your
- 6:46brand has a million people purchasing
- 6:48products from it, let's say you're
- 6:49Alibaba or Amazon, it just doesn't make
- 6:52sense and it's very expensive. So from a
- 6:53harness perspective, you want to be
- 6:55smart about this and you want the system
- 6:56to be automatic. And then a simple way
- 6:58is probably like maybe consolidate these
- 7:00timeordered events after every say 2,000
- 7:02conversations because you have a million
- 7:03customers and then you can feed these
- 7:05things into a summarizer agent which is
- 7:07another large language model harness.
- 7:09You can define the system prompt in this
- 7:10one. You can probably feed it with some
- 7:12memories too. Uh you can configure
- 7:14different models. Maybe it could be a
- 7:16cheaper model because you're feeding too
- 7:18much text into it. So the context window
- 7:20is too big and probably these are very
- 7:22expensive. So you can use cheaper open
- 7:24source models if you want to. Having
- 7:25such a mechanism allows you to
- 7:28consistently update this memory system.
- 7:31The data should be coming from the
- 7:34previous large language model replies.
- 7:36Again, let's review how this harness
- 7:38works. A user sent a prompt. We're in
- 7:40one agent runtime with the current chat
- 7:42history and how the agent should be
- 7:44performing the system prompt. We're
- 7:46preparing a working memory for this AI
- 7:48agent to be able to answer a question.
- 7:51And after every single time it answered
- 7:52a question, it will send these messages
- 7:55to this database. And then this database
- 7:58is basically feeding back to this
- 8:00working memory every single time when a
- 8:02question is checking for relevant
- 8:04context. And at the same time, because
- 8:06this database is too big, sometimes you
- 8:09want to consolidate them into some
- 8:11summarized information or distilled
- 8:13facts so that they're stored properly in
- 8:16a semantic memory so that the retrieval
- 8:18of such memories is just faster. I know
- 8:20we talked about retrieval a lot and
- 8:22that's just another buzz word called rag
- 8:24which is retrieval augmented
- 8:25generations. I also have a few videos
- 8:27explaining what rags are. Feel free to
- 8:29watch them. There's a little bit of
- 8:30difference between how retrieve from
- 8:32semantic memory and episodic memory. For
- 8:34semantic memory it's just racks because
- 8:36these are just facts and text or files
- 8:39right but then for episodic memory
- 8:41remember this is a time series. Let's
- 8:42say we're still in this e-commerce store
- 8:43right the user question could be like
- 8:45what were the previous 10 conversations
- 8:46that we had with this specific customer
- 8:48from the United. And then you might just
- 8:50need a SQL query to query something
- 8:51that's pretty recent from this episodic
- 8:54memory. But if your question is like
- 8:56what were my previous 20 conversations
- 8:58that have customer complaints on the
- 9:00quality of the products and our agent
- 9:02did not successfully resolve with such a
- 9:05question you not only need a SQL query
- 9:07which is just capturing the data events
- 9:10in a data table. You also want to do
- 9:13some semantic search and that's why here
- 9:15rag is important because it's checking
- 9:17for relevant information for you. You
- 9:19don't want the entire 2,000 messages.
- 9:21You want that 20 messages out of these
- 9:232,000 that are exactly relevant to what
- 9:25you want. And because these complaints
- 9:27are in text, we need to do some
- 9:29retrieval augmented generation to match
- 9:32the semantic meanings between text and
- 9:34the user prompt so that you're fetching
- 9:36the right context for the working
- 9:37memory. By now, this is probably a fast
- 9:40walkthrough of the memory system again,
- 9:42but we're just thinking about it from a
- 9:44hardness perspective in this video. And
- 9:46remember, for hardness, we're training
- 9:48this horse of LLM to run autonomously
- 9:50without having too much randomness.
- 9:52Okay, there's another piece of it that's
- 9:54quite important, which is the agent
- 9:55might not only just read the memory. It
- 9:58might also do some tasks or call some
- 10:01tools. When an agent calls tools, it
- 10:04might not necessarily be just one time
- 10:06call. It could be multiple times of
- 10:07calls. For instance, let's say this AI
- 10:10agent has a bunch of agentic tools such
- 10:13as help me schedule a meeting, help me
- 10:15read or write my customer relationship
- 10:17data from the CRM system, or help me
- 10:20fetch the payment information, say from
- 10:22Stripe or Alip Pay. And here's something
- 10:23we should be careful about. If we give
- 10:25this horse or this LLM technology full
- 10:28power, it could just continuously do
- 10:30this forever, right? or it might not
- 10:32even know what's the right time to stop
- 10:34or what is the right tool calls it
- 10:36should make when is the endpoint to
- 10:38decide okay this response is good enough
- 10:40let's move on to reply that's why we
- 10:42have this mechanism called end loop
- 10:44guardrails yes now we're talking about
- 10:45loop engineering one of the biggest
- 10:47buzzwords in the recent few months a
- 10:49loop is part of harness why because a
- 10:52loop is also helping us to control this
- 10:56technology to make sure it runs the way
- 10:58we want it to run and example that could
- 11:00helpful for you is that let's say the
- 11:03custom prompt is help me find out what
- 11:06customers are complaining about our
- 11:08products. What are some of the
- 11:09follow-ups we could do in order to uh
- 11:12win them back? And if they're asking for
- 11:15reimbursement, have we done the
- 11:16reimbursement or not? If not, can we do
- 11:18that? This is probably a series of
- 11:20questions, but sometimes we just dump
- 11:22all of these things into AI agent. Okay.
- 11:24And after you have this prompt, this LM
- 11:26Asian needs to decide, okay, what are
- 11:29some of the tools that could be helpful
- 11:31for me to finish this task? Loop here is
- 11:33basically an architectural thinking of
- 11:36when is good enough so that we stop and
- 11:39give the user or the business owner a
- 11:42reply. Okay, so what might happen here
- 11:44is that the LM agent is doing a bunch of
- 11:46tool calls. It's doing some thinking.
- 11:48It's saying let me read from our
- 11:50customer relationship management tool
- 11:52like Salesforce, HubSpot or Automatus
- 11:54and then it's going to find out okay
- 11:56there were 30 customer complaints in the
- 11:59past 2 months 12 of them have got
- 12:01reimbursement the other eight have not
- 12:02got reimbursement. So after the first
- 12:04initial fetch, it's probably responding
- 12:06to the AI agent, right? And the agent
- 12:08will be probably thinking, okay, the the
- 12:10task or the ask is that um can we can we
- 12:13follow up with some of those who did not
- 12:14get um the reimbursement, right? So and
- 12:17then they probably like just make
- 12:18another tool calls and be like, hey,
- 12:20let's schedule a meeting with those
- 12:23customers who did not get a
- 12:24reimbursement, which are the eight of
- 12:26them. If we go a little bit more
- 12:27advanced, we even just use the
- 12:28reimbursement trigger on Stripe or Alip
- 12:31Pay to refund the customer. Can you see
- 12:32that this is a loop until we finish the
- 12:34task? But of course, this is like a case
- 12:36by case situation. It really depends on
- 12:38what your task is, how you built the
- 12:40system. So, there's no one solution fits
- 12:42all here. I'm just explaining what a
- 12:44loop is. The very very essential part of
- 12:46this loop is that it needs to know when
- 12:48it should stop. That's why we need this
- 12:50end loop guardrails. The guardrails
- 12:52could just simply be the task is done
- 12:54and perhaps when the agent was doing the
- 12:55planning, it should confirm with the
- 12:57user what is a good ending point. It
- 12:58might clarify with you, is this what you
- 13:00want? reimbursing the other eight people
- 13:02or should I just tell you who they are
- 13:03and then you will follow up later.
- 13:04Right? These are two different decisions
- 13:06you can make and after you make you're
- 13:08basically telling the agent loop that
- 13:10there's an ending scenario. Another good
- 13:12example I saw today is that you know
- 13:14when you're doing coding and claw code
- 13:16can just always pop up some windows and
- 13:18ask you for permissions right so the
- 13:20good way to use a loop engineering here
- 13:22is that you can set up a loop or set up
- 13:24a hook in claw code and telling it that
- 13:27you should always send me a notification
- 13:28on my laptop if you are pending on some
- 13:31permissions from me otherwise if I'm
- 13:32watching YouTube and then when I come
- 13:34back 30 minutes later I realize that
- 13:36clock code is stuck in that one
- 13:38permission like 25 minutes ago that
- 13:40would be a waste of my time. Okay. So,
- 13:42you can set up a loop like this to make
- 13:44sure that there's a way to send you
- 13:46notification pop-ups so that you know
- 13:47the loop has ended or it needs your
- 13:49input again. Are you guys still with me?
- 13:51Good. So, by far we have covered AI
- 13:54agent run with a memory system with a
- 13:57loop engineering around the large
- 13:58language model agent which has a trigger
- 14:01to end the loop so that it sends a reply
- 14:03to the user and basically this whole
- 14:05thing is an AI agent harness system.
- 14:07What's next is one of the other biggest
- 14:10buzzwords that Y combinator always
- 14:12mentions which is aval or LLM ops. Let's
- 14:15jump into it. But firstly I want you to
- 14:17understand why do we need LM ops here.
- 14:20Let's still look at the left hand side
- 14:21with this hardened system. The biggest
- 14:23problem here is that we don't know how
- 14:25well it's performing and that's why we
- 14:27need a feedback loop to help us
- 14:29understand is this agent actually
- 14:32performing properly right for my
- 14:34business or for my use case. And can I
- 14:37continuously get feedback on how do we
- 14:40fix it and actually fix it ourselves?
- 14:42Okay. And when we say fix it, a simple
- 14:45way to understand it is that can we have
- 14:47a better system prompt? Can we have a
- 14:50better large language model
- 14:51configurations?
- 14:52Is there something we should change for
- 14:54how we retrieve the AI agent memories?
- 14:57These are kind of things that we can
- 14:58continue to iterate. But in order to
- 15:00iterate to make sure this system runs
- 15:03properly, we need a way to evaluate it,
- 15:06diagnose problems, solve the problems
- 15:09until it's a healthy and wellperforming
- 15:12system. And that is called large
- 15:14language model operations system, LLM
- 15:17ops. So again, in order to understand
- 15:19this properly, we need to come back to
- 15:22what an agent run is. So an agent run
- 15:24you can simply understand it as a user
- 15:26question is sent to a large language
- 15:28model and they get a reply that is one
- 15:31agent run but in this agent run the
- 15:33agent tool calling could happen multiple
- 15:35times that does not matter right we're
- 15:37just talking about from a user input to
- 15:40a response from agent perspective that's
- 15:42one agent run and then we're going to
- 15:44introduce this system called a tracing
- 15:46system so every agent run we should
- 15:49trace like a tree of events that
- 15:52happened and there are lots tools that
- 15:53could help you with that. It could be
- 15:54lenuse, could be lens, etc., etc. A tree
- 15:56of events could be like what did the
- 15:58person actually ask? What retrievalss
- 16:00did the model actually retrieve? How
- 16:02many times did the large language model
- 16:03actually call the tools and how was the
- 16:05tool usage, how was the response time,
- 16:08right? How long did it take for this
- 16:10entire system to run for checking
- 16:11latencies and how many tokens have we
- 16:14used when we do these tool calls, agent
- 16:17run, you know, doing this retrieval,
- 16:19augmented generation, these kind of
- 16:20things. So trace is helping us to track
- 16:23events basically and that's the first
- 16:25step. This is the step first to collect
- 16:26data and these data will be used for the
- 16:30following two purposes. Was it a good
- 16:32system run and was it healthy which
- 16:34corresponds to evaluation system. We can
- 16:37probably use large language model as a
- 16:39judge here to give us a score on how
- 16:42well it performed. For example, if the
- 16:44task had something to do with scheduling
- 16:47meetings, did the meeting actually
- 16:48triggered? How long was the response for
- 16:50an agent to reply to a question? Was it
- 16:5220 seconds or was it two milliseconds?
- 16:54And also things like how many tokens
- 16:56have we used? These two are basically in
- 16:58the same system. You can write it as
- 16:59deterministic code. You can use an AI
- 17:01agent to do it. But this is like part of
- 17:03the procedure which is helping us to
- 17:04understand was this a healthy system and
- 17:07was it a good system. And after that
- 17:09we're going to diagnose okay where and
- 17:12why something was broken. For example,
- 17:14the meeting scheduling event was never
- 17:16triggered. Why was that? Right? Okay, we
- 17:18want to understand why was that and we
- 17:20could probably feed that into a coding
- 17:21agent in claw to sort of deep dive into
- 17:23it. Or you know if the latency is 20
- 17:26second instead of 2 milliseconds
- 17:28something's wrong. Maybe one of the tool
- 17:29call is taking too much time. Maybe the
- 17:32working memory is too large. Uh so that
- 17:36the response time for a large language
- 17:37model to a memory retrieval is just
- 17:39taking too much time. Maybe not every
- 17:41single question requires a retrieval
- 17:43from all these gigantic memory system.
- 17:46Maybe you're just asking a simple
- 17:48question be like when was my birthday?
- 17:49When was open AI started and these kind
- 17:51of information you probably don't need
- 17:52to do a ton of retrieval. The model
- 17:54itself already knows. So you basically
- 17:56want this system to provide a dashboard
- 17:58for you to understand the metrics and
- 18:00then with these metrics you can diagnose
- 18:02what is going wrong. And then we're
- 18:03going to have a little gate here which
- 18:05is if the evaluation system passed well
- 18:07you can define the rules. We can either
- 18:09ship some very simple fix have a new
- 18:11version of the prompt or update the
- 18:13model configuration you know some tool
- 18:15changes or the parameters for
- 18:17retrievalss the LM ops will feed the
- 18:20improved system prompt and the
- 18:22configuration of the model back to this
- 18:24agent run system when then one LM ops
- 18:26loop is finished if let's say something
- 18:29is deeply broken right we cannot just
- 18:31simply ship the latest version of the
- 18:33prompt then we should go fix the bug
- 18:36rerun the agent run resend the question
- 18:38and and then retrace the events and then
- 18:40redo this evaluation system in this LLM
- 18:43ops architecture. So now let's zoom out
- 18:45and look at this chart one more time. We
- 18:48covered what an AI agent run is. We
- 18:50covered how it would retrieve
- 18:52information from memories and we
- 18:54understood how an LLM agent would ask
- 18:57questions would call tools to help it
- 18:59finish the task in a loop and it knows
- 19:01when to stop the loop so that we can get
- 19:03the reply. And this whole thing is a set
- 19:05of harnessed tools that we're
- 19:06controlling this horse, this technology
- 19:09to run in the right direction. Okay, to
- 19:11do the right task and at the same time
- 19:13we have like a health checking system or
- 19:15evaluation system to understand how
- 19:18every single run is being traced, is
- 19:20being observed and how do we diagnose
- 19:23some problems and fix some problems and
- 19:25ship the latest updates of the prompt
- 19:28about the model configuration about all
- 19:30these parameters or knobs that needs to
- 19:32be updated. so that this system will be
- 19:35an autonomous system that will just
- 19:37self-evolve and grow over time. I really
- 19:39hope this was helpful. Let me know what
- 19:41you think and you have any questions,
- 19:42you can always reach out to me. I'll see
- 19:44you in the next video. Thanks so much.
- 19:45Hi everyone, this is Sean. So today
- 19:47we're going to walk through Hermas agent
- 19:49harness and it loop engineering system.
- 19:50This is one of the most popular harness
- 19:52agent system right now and it's an open
- 19:54source project. If you check their
- 19:56GitHub, they've got more than 200,000
- 19:58stars in a very short period of time.
- 20:00So, it's basically a self-improving AI
- 20:02agent built by this research lab called
- 20:04new research and it's a built-in
- 20:05learning loop. It creates its own skills
- 20:07from experience instead of you telling
- 20:09the AI like claw code to make skills and
- 20:11then it also improves over time when you
- 20:13keep using it because it persistently
- 20:15stores knowledge in your local machine
- 20:17like a MacBook. But anyways, this is all
- 20:18the buzzwords. We're going to jump into
- 20:20this Hermes Asian system design like
- 20:22usual. We're going to test it as well
- 20:24using their desktop app. Just give you a
- 20:26quick example. It's like a claw code,
- 20:27right? You can see that it was thinking
- 20:29and then it answered me in like a
- 20:30Pikachu style. It said pika pi just hang
- 20:32out in the my directory because it's
- 20:34literally a bot living in my local
- 20:35machine files. What we're building
- 20:37fixing today pika and also you can
- 20:39interact with it on WhatsApp and then
- 20:40just say good morning. It's a little bit
- 20:43slow. I'm using a Gemini 3.1 model for
- 20:46this. It took like a good 5 seconds and
- 20:48now it's basically just using WhatsApp
- 20:50as the gateway to receive information
- 20:52from my WhatsApp account. So I can just
- 20:54text it and then ask it to do things for
- 20:56me and say picapy good morning. You
- 20:57probably realize that I've changed his
- 20:59personality a little bit and there are
- 21:00more interesting things I can show you.
- 21:02For those who have watched my previous
- 21:03agent harness lube engineering and
- 21:05memory system you might realize that
- 21:06this chart is somewhat similar. This is
- 21:08what the previous chart looks like. And
- 21:10now if we look at harness for Hermes
- 21:12agent again I added a bit more
- 21:13information especially the difference
- 21:14between the herist harness versus a
- 21:17generic harness system. I marked the
- 21:19differences in red but we're going to
- 21:21walk through this step by step. Also, I
- 21:22prepared some testing examples to cover
- 21:24harness, memory, skills, loop, gateway,
- 21:28and LM ops or eval because a lot of
- 21:30people were asking me if I could
- 21:31implement some real examples. So, here
- 21:32we are. Without further ado, let's jump
- 21:34right in. Okay. So, firstly, let's look
- 21:36at the foundation of Hermes harness
- 21:38agent. It's a harness that runs on your
- 21:40local machine. You can use either local
- 21:42command lines, docker, ssh or a virtual
- 21:45VPS. There are two main ways to interact
- 21:46with Hermes agent. One is using your
- 21:48favorite communication apps. And in my
- 21:50case, I used WhatsApp or you can just
- 21:52use their desktop app with the Hermes
- 21:54agent. So you can come to Hermes
- 21:55official website
- 21:56herm-aggent.newresarch.com
- 21:59and just click on download for Mac OS or
- 22:01you can click on Windows over here to
- 22:02use the command line to install it.
- 22:04Essentially still a chatbot just like
- 22:06claw code and you ask the question to
- 22:07either the communication app or the
- 22:10desktop version. It's very similar to
- 22:11open claw where you can use WhatsApp to
- 22:13control it. So after user send the
- 22:15prompt, the user prompt with the current
- 22:17chat history and the system prompt will
- 22:19be fed into a working memory. And then
- 22:21this LM agent is basically the Hermes
- 22:22agent that will be answering your
- 22:24question and it will be using a loop
- 22:26engineering over here to call up some
- 22:28tools that are only on Hermas. And
- 22:30eventually after finish the task, it's
- 22:32going to end the loop and then send the
- 22:33user a reply. Just literally sending a
- 22:35question throughout this entire firmal
- 22:37agent run and then get the user reply
- 22:39just like my Hermes was greeting me like
- 22:41a Pikachu. And by the way, how did this
- 22:42happen? Right, I'll show you right now.
- 22:44So, the system prompt for Hermes is
- 22:46interesting. It's called soul.md. I
- 22:48really like how they name it. It says
- 22:50soul. If you come to the right hand
- 22:52side, you can see there's a little file
- 22:53bar and this is your local files and you
- 22:56can go find out there's a hermas folder
- 22:59and we open that. And if we scroll down,
- 23:02you will see there's a file called
- 23:04soul.md. Double click on that. And you
- 23:07can edit this yourself, right? But since
- 23:09I've already edited it, after you edit
- 23:10anything, you can just click on save. I
- 23:12basically say you talk like Pikachu.
- 23:14Every reply starts and end with pika
- 23:16pika pikap. If you're excited, say pikap
- 23:18with the lightning. If you're sad, say
- 23:20quiet pika. I'm pretty sure you can do
- 23:21something similar in claw code for this.
- 23:23In claw code, something similar will be
- 23:25in customize general. And you can see
- 23:28there's an instruction for claude. You
- 23:30can tell claw to do certain things and
- 23:32not to do certain things. It's basically
- 23:34the system prompt used across the entire
- 23:36agent harness. If you didn't watch my
- 23:37previous video on harness and loop
- 23:39engineering, harness is basically a set
- 23:40of tools that allows you to control over
- 23:43this really powerful horse which is the
- 23:45LLM itself. In our case, the LM we're
- 23:48using is Gemini because I've got the
- 23:50Google credits. I don't want to burn my
- 23:51anthropic credits, but feel free to use
- 23:52anthropic. This is the this is the agent
- 23:55of the the horse and the harness is this
- 23:57entire thing. So it's like a horse
- 23:59harness that you ride on and you make
- 24:01sure that the horse is not going
- 24:02anywhere random and it's using the tools
- 24:05accordingly to move in the right
- 24:06direction. Harden is a very vague
- 24:08concept. Just don't fantasize it. It's
- 24:10nothing complicated. It's just a buzz
- 24:11word. Then we're going to continue to
- 24:12build up this harness right here for
- 24:14Hermes. But what I'm going to show you
- 24:15is what a loop is. The concept loop got
- 24:18really popular because these days people
- 24:20build a lot of agents. You don't want to
- 24:22always tell the agent what to do. You
- 24:23want the agent to figure out what kind
- 24:24of prom you should send, what kind of
- 24:25tools you should call. Let's take a look
- 24:27at what tools Hermes agent actually have
- 24:30to enable this loop engineering. We're
- 24:32going to cover this briefly and then
- 24:33I'll show you real examples. So first of
- 24:34all, the agent of Hermas is having a set
- 24:37of tools that include the terminal, the
- 24:41browser. Okay, you can use your laptop's
- 24:43terminal, you can start a browser, it
- 24:45can delegate task, which means that it
- 24:48can spawn up some sub agents for you.
- 24:49For example, you can ask it to spawn up
- 24:51an agent which will call your claw code
- 24:53CLI to write code for you if the task is
- 24:56basically fixing some bugs on my GitHub.
- 24:58And then it can schedule a chrome job. A
- 25:00chron job is basically it can schedule
- 25:01something that will happen at certain
- 25:03time. You don't need human intervention
- 25:04to do it. For chronop, let's just try
- 25:06right now in Hermes Asia. I can say help
- 25:08me set up a chrome job that will tell me
- 25:09a Pokemon joke every minute in the next
- 25:1110 minutes. Okay, I'm just going to send
- 25:14this and also I'm going to send this to
- 25:16our Hermes Asian WhatsApp interface,
- 25:18too. And instead I'll just say tell me a
- 25:21developer joke every it just quickly
- 25:23called this a chrome job and it says it
- 25:26repeats 10 times for the next 10 minutes
- 25:28because it says it's hanging in the
- 25:29terminal UI it won't be able to print it
- 25:31so I need to ask it I just ask like
- 25:34what's the first joke and you can see on
- 25:36WhatsApp it also got scheduled the
- 25:38chrome it say the joke machine is
- 25:40powered and scheduled I've set up a
- 25:42chrome job to send you a fresh joke
- 25:44every minute in the next 10 minutes you
- 25:46should receive the first one in about a
- 25:47minute we'll see about that Okay. And
- 25:50why did the Squirtle cross the ocean to
- 25:52get the other tide? Pika pika. Okay. So
- 25:54that was Chrome job. It can do skill
- 25:56management as well. It can connect to
- 25:57MCPS as well. It depends on the question
- 25:59that the user asks. The agent will just
- 26:01leverage these tools until it's done and
- 26:04then it will send a reply to you. Okay.
- 26:06The example we just saw was Chrome job.
- 26:08I'm going to show you some more examples
- 26:09later too. Let's test a few hardness and
- 26:11loop examples. The first one is use our
- 26:14terminal tool to find out what OS
- 26:15version I'm on and read the last 10
- 26:17lines of my shell history. Let's do
- 26:18that. Start a new session. I'm just
- 26:21going to ask this exact question to it.
- 26:23So you can see it ran this on my local
- 26:25machine and then it figure out that I'm
- 26:27on Mac OS 26.5. And the last 10 lines I
- 26:30used in shell history is basically text
- 26:33get status get status get status. I
- 26:35always like to check get status. And
- 26:36here what happened was that it ran
- 26:38through this agent run and then it
- 26:40called the terminal to do the task for
- 26:43me which is checking what system I'm on.
- 26:46This is a very simple example but you
- 26:48can do more crazy things. You can ask it
- 26:49to use the local machine terminal to do
- 26:51things for you. If you're using clock
- 26:53you're you're very familiar with this
- 26:54already, right? There's nothing special
- 26:55about it. I'm just showing you the
- 26:56example that this is a harness and
- 26:58there's a loop because after I fetched
- 27:00it, it will stop and tell me and reply
- 27:02to me. Let's look at the chron job as
- 27:03well. You can see they sent me two jokes
- 27:05already. Why do Java developers wear
- 27:08glasses? Because they don't car. Pika
- 27:11pika. I would tell you a UDP joke, but I
- 27:14have absolutely no guarantee that you
- 27:16would get it. I'm not going to wait for
- 27:17acknowledgement anyways. Pika. Okay,
- 27:20this one I didn't get it. So, one more
- 27:22example. Browse the YouTube channel
- 27:24Sean's AI stories. Title plus views. My
- 27:26last videos. Create new video titles for
- 27:29this current video I am filming. So,
- 27:32here we're asking you to use the
- 27:33browser. basically okay you can see that
- 27:35it started doing the searching all these
- 27:37things already exist on claw code I'm
- 27:39just explaining it from a system design
- 27:40perspective and showing you examples
- 27:41okay and then also I will show you
- 27:43what's the difference between her and
- 27:45claw code so please bear with me all
- 27:47right so I need to run a python script
- 27:50so you can see also use some clicking
- 27:52tools oh there's an error the website
- 27:53was wrong
- 27:56try this new one
- 27:58it saved a little memory here okay it
- 28:01said user profile updated. This is
- 28:04something I'm going to cover in a bit.
- 28:06But spoiler alert, Hermes As a Asian
- 28:08will save agent memories by itself. It's
- 28:11basically in your local files too. If
- 28:13you come here within the Hermes folder,
- 28:15if you scroll down in memories, you can
- 28:17see there's a memory. MD. If we double
- 28:19click on that, it stored YouTube here.
- 28:22YouTube scraping quirk. So, because it
- 28:24realized that it had a mistake, so it
- 28:26updated itself with the memory. This is
- 28:28a self iterating process of this
- 28:29harness. Good. So, it read my previous
- 28:31YouTube titles. it realized that I've
- 28:33got this video that's got 100,000 views
- 28:36and since I'm filming a new video now it
- 28:38gave me some new titles. So what
- 28:39happened here was the user which is me
- 28:41who asked the question through the
- 28:43desktop and then it went through this
- 28:45agent run and then this Asian harness
- 28:48Gemini that I'm using was basically use
- 28:50calling the browser right here to search
- 28:52up my YouTube channel and then it came
- 28:55back and said hey I didn't find
- 28:56anything. So uh it basically stopped.
- 28:59There's an anal loop guardrail. The
- 29:01guardrail stopped and say it didn't find
- 29:02anything. So I realized that okay the
- 29:04previous URL was wrong. So I input it
- 29:06again. Uh it called the browser again
- 29:08and it also called the terminal run some
- 29:10Python script because you needed to do
- 29:12some task and eventually it came back
- 29:14with the answer. There's a mechanism
- 29:16that says okay end the loop. No more
- 29:18tool calls and send the answer. These
- 29:20are two quick examples of what a hard
- 29:22loop engineering is. You see you're
- 29:23already using it every single day. It's
- 29:25nothing fancy. So don't get overwhelmed
- 29:26next time when you hear these terms.
- 29:27It's just a buzz word. So what happened
- 29:29just now with this memory MD that we saw
- 29:32regarding YouTube memory Hermes agent
- 29:35after or during the agent run, it's
- 29:37going to start to create its own
- 29:39skill.md and memory MD depending on if
- 29:42it's feeding into the procedure memory
- 29:45or if it's feeding into the semantic
- 29:47memory. And we should add one more arrow
- 29:48here basically. So essentially the way
- 29:50that Hermes agent works is that it's
- 29:53storing its own skills and memories
- 29:55completely locally. Nothing is stored on
- 29:57the cloud. It has its own self-improving
- 29:58loop. Every time when it realize that
- 30:00something is a mistake they should learn
- 30:02from or it's a repeated task, it will
- 30:04start to summarize. Okay, here are some
- 30:06new memories I should know for the
- 30:08future in order to serve this user well.
- 30:11It started to save these files into
- 30:14these two chunks of memory system. First
- 30:16one is called procedure memory. is
- 30:17basically how you act, how the horse
- 30:19agent should be acting and it's usually
- 30:21saved in this directory called Hermas
- 30:24skills skill.md her skills and any of
- 30:27these okay autonomous AI agent claw code
- 30:31there's a skill for how Hermes will use
- 30:33claw code here so it will basically
- 30:35delegate any coding to claw code shows
- 30:38exactly how you should do it like
- 30:39install claude run it and do the
- 30:42authentication and all these kind of
- 30:43things and then what we just saw earlier
- 30:45was memories memory saving the semantic
- 30:48memory which are some durable facts or
- 30:50things related to the user profile. So
- 30:52if the user has a habit or has some
- 30:54important facts or information he should
- 30:56remember then it will be saved in the
- 30:57memory immd. And what's interesting is
- 30:59that for Hermas instead of doing the
- 31:01embeddings it's actually just using
- 31:03plain text. Okay, it's using the top
- 31:05cake keyword instead of doing the
- 31:06embedding or the rack system. It's very
- 31:08interesting because I'm not too sure why
- 31:10but it's just doing text. So remember
- 31:12that we say the agent run is ephemeral.
- 31:14Now that with a working memory to be fed
- 31:17into the agent for all the loop
- 31:18engineering which includes procedure
- 31:20memory, semantic memory and later we'll
- 31:23cover episodic memory. This will become
- 31:25a more complete memory system that will
- 31:28feed into the agent runs which is making
- 31:30the harness more elegant. And the third
- 31:32pillar that we haven't covered is called
- 31:33episodic memory. It's basically the chat
- 31:35history or some data events. It can be
- 31:37ragged or sequed into the working
- 31:39memory. After every agent run, the
- 31:41information will flow into the episodic
- 31:44memory. What's different in Hermas is
- 31:46that the data is not saved in the cloud.
- 31:48It's saved in this place called
- 31:50state.db. Again, let's come back here on
- 31:52the right hand side bar. If we again
- 31:54open her folder, scroll down. You can
- 31:57see at the end there's a state db. I can
- 32:00click on preview. Anyway, this is
- 32:01basically a database that will include
- 32:04all those info in your chat history, but
- 32:06also over time it will start to
- 32:08consolidate some of those chats using
- 32:10the Heras auxiliary models which are
- 32:12cheaper non-important models to do this
- 32:14summarization task so that it will
- 32:16distill some facts into what we call the
- 32:18semantic memory or that memory MD. So,
- 32:20you can see this harness is running
- 32:23really in a self-improving autonomous
- 32:26situation which is really cool. I know
- 32:28this chart is getting more complicated
- 32:29than usual, but I feel like this should
- 32:31be a complete summary for how the Hermes
- 32:33agent works. And let's go through a few
- 32:34more examples. Let's test out some of
- 32:36the stuff for memories. Save to memory
- 32:38that my favorite testing framework is
- 32:39piest. So this is an explicit saving,
- 32:42right? It updated the memory file. So if
- 32:45we go to memory,
- 32:46you can see one more block which is
- 32:48user's favorite testing framework is
- 32:49piest. And that was an explicit ask from
- 32:52the gateway. The tool is basically let's
- 32:54go update the memory over here. Search
- 32:56our past sessions. What was the very
- 32:58first thing I ever said to you? Now it's
- 33:00using the searching tool in the loop
- 33:02engineering iteration. Said the first
- 33:04thing I said is what can you do Hermas
- 33:06and what makes you special? Let's double
- 33:08check that. That's right. And it failed
- 33:10because I couldn't log in at that time.
- 33:12So what happened was I asked the
- 33:13question through the interface. It went
- 33:16through the prompt and in the working
- 33:17memory because we asked that. So it
- 33:19curried the information from episodic
- 33:21memory from state. Db. It returned the
- 33:24memory into the agent. And the agent was
- 33:26like, "Oh, okay. This is what you said."
- 33:27Let's try an example related to skills,
- 33:29which is procedural memory. So, we're
- 33:30going to say, "Create a skill called
- 33:32video prep that captures how I format my
- 33:34video scripts. Spoken English, define
- 33:35jargon on inline, no m dashes, closes
- 33:38with you can build anything, you can
- 33:40learn anything." This is an explicit
- 33:41call to update the skills. It used the
- 33:44skill management tool. Remember in the
- 33:46loop engineering, there's a tool called
- 33:47skill management. And this is what it
- 33:48did. And say, "Picape, I have
- 33:50successfully created the video prep
- 33:51skill. Let's find out about it." again.
- 33:53Hermes skills media video prep. Double
- 33:58click. It created this skill for me. It
- 34:00says no m dashes spoken English define
- 34:03jargon in line catchphrase. You can
- 34:05build anything. You can learn anything.
- 34:06Good. So we can test the loop a little
- 34:08bit more. We can say use delegate task
- 34:10to spawn two sub agents. One is
- 34:12researching the LM eval harness. Another
- 34:15one researching VLM architecture. The
- 34:18agents was calling the tool caught
- 34:19delegate task and successfully spawned
- 34:21two background agents. So you can see
- 34:23it's processing right now over here.
- 34:25They're working in parallel.
- 34:27Cool. Eventually return the results to
- 34:30me. So you can see that it's spawn up
- 34:31some sub agents to do task for me in
- 34:33parallel. We've already tried the chrome
- 34:35job. And the next one is spawn sub agent
- 34:38that uses claw CLI. This is my favorite.
- 34:40So this is my favorite. Spawn a sub
- 34:42agent that uses the claw CLI in headless
- 34:44mode. basically using claude to build a
- 34:46Python script fetching the top five
- 34:48hacker news stories to markdown and
- 34:50verify it runs. Yep, it firstly
- 34:52delegated a task to claude CLI is using
- 34:56the skill. You can see remember we read
- 34:58the skill here in autonomous AI agent
- 35:00clot. It should be reading this skill.
- 35:02Is this helpful? Let me know if this is
- 35:03helpful. I mean, I was trying to show
- 35:05you more examples in the walkthrough so
- 35:07that it feels more concrete because a
- 35:09lot of people were asking me for
- 35:10implementations from my previous videos
- 35:12and I feel like Hermas is really good
- 35:14example.
- 35:17Okay, it created this uh code for me and
- 35:20used the cloud. It created the fetch
- 35:23hack news.py and just say show me the
- 35:27result after
- 35:30running the script.
- 35:34Yep. So these are the top five hacker
- 35:36news stories on hacker news. Let's come
- 35:38check it's correct. Remember what
- 35:39happened was that we sent the question
- 35:41from the gateway. It went through the
- 35:42agent run and the LLM basically caught
- 35:45up the tool called delegate task.
- 35:47Delegate task delegated the task to sub
- 35:49agent and the sub agent was running claw
- 35:51code and cl code wrote that script and
- 35:53sent it back and then this agent run the
- 35:56script again and then we get the
- 35:58results. This was the harness in the
- 35:59loop engineering in this entire thing. I
- 36:01don't think it self-t triggered
- 36:03self-learning skills yet. It only
- 36:06triggered learning the memory. I'm just
- 36:08going to ask explicitly. Did you save
- 36:09any skills by yourself or did you only
- 36:11save memory so far? Oh yeah, but we
- 36:13created that. But we created
- 36:16that skill. You did not self realize
- 36:21you need to save skill. You should
- 36:25proactively offer to save a skill after
- 36:27difficult iterative task with five or
- 36:29more tool calls. But so far we haven't
- 36:31done anything difficult. So we can test
- 36:32the gateway. I can summarize what we
- 36:35worked on today. H this chrome job
- 36:37failed. It didn't tell me 10 jokes. It
- 36:39only told me two. What were the rest of
- 36:43the jokes? Don't lie to me because maybe
- 36:46you just created it. Anyways, Hermes
- 36:48team, if you're watching this, your
- 36:49chrome job is not updating it unless I
- 36:53ask for the results. So maybe this is
- 36:55something you guys can fix. All right,
- 36:56let's move on and summarize what we
- 36:58worked on today because in this current
- 37:00chat, it doesn't know the rest of the
- 37:03stuff here. So, let's see if it's
- 37:04calling the actual remember we say state
- 37:07DB, memory MD, and skills.m MD. Good.
- 37:10Actually knows. Okay, it says jokes on
- 37:12demand system check on the OS version,
- 37:15which is what we did. YouTube stats.
- 37:17Yes, we did that. Hacker news. Okay,
- 37:19good. WhatsApp here is just a gateway as
- 37:22an interface. It's doing the same thing
- 37:24as the Hermist agent desktop. This is
- 37:26very cool. I mean, this is similar to
- 37:28OpenC Claw, but still every time when I
- 37:30feel I can just have a personal
- 37:31assistant handy on my WhatsApp. This is
- 37:33really cool. And imagine if I host this
- 37:35on a virtual machine and it runs
- 37:37forever. I can just basically ask
- 37:39WhatsApp what's going on and it's going
- 37:41to reply to me. Ooh, I completely
- 37:42forgot. You should set up WhatsApp by
- 37:44typing in her WhatsApp and then just
- 37:46follow the instructions to set it up.
- 37:48Should be straightforward. Last but not
- 37:50least, where is Eval? All right, where
- 37:52is our LM ops? Remember in our original
- 37:55chart we had this blue box on the right
- 37:56hand side called LM ops. So
- 37:59unfortunately on Hermas based on my
- 38:01current research I don't think it has an
- 38:03LM ops or an Eval system. You probably
- 38:06need to build it yourself but it does
- 38:08trace the run. There's a thing called
- 38:10trajectory export and logs and it's
- 38:12basically just logging the whole thing
- 38:13but there's no eval. If you watched our
- 38:15previous video, the eval basically you
- 38:16can use tools like land smith land view
- 38:19fuse and all these things that can track
- 38:21the entire agent run and some events
- 38:23that happen like tool calls. How many
- 38:25times did you call LLMs? Did you use
- 38:27some cheaper models to summarize things
- 38:29into your semantic memory?
- 38:30Unfortunately, Hermas doesn't do that.
- 38:32I'm curious why. Maybe because it's a
- 38:34local system, so you can just customize
- 38:36it yourself. It's under Hermas. You can
- 38:38click into logs. You can see all the
- 38:40logs here. Can see the errors. can see
- 38:42the agent runs. Gateway logs to gateway
- 38:45is the entry point which is WhatsApp or
- 38:47this desktop. I hope this was helpful. I
- 38:49think there's quite a huge amount of
- 38:52work that we covered today which are
- 38:54exactly what claw code can already do or
- 38:56open claw can already do. I feel that
- 38:58what makes her special at least compared
- 39:00to claw code is that it's self updating
- 39:03all these things locally. And maybe this
- 39:05is good for privacy reasons because I
- 39:07don't see why we should not save it on
- 39:10the cloud. Other than that, everything
- 39:11else should be very similar, right? It
- 39:12runs the loop every time when the agent
- 39:14needs to do some task uh kind of call
- 39:15all these tools and you can even
- 39:17schedule chrome drop and then the
- 39:18memories and the skills are getting auto
- 39:21updated uh into the procedure memory,
- 39:23semantic memory and the state data will
- 39:25be saved into the episodic memory. So I
- 39:27would say this is a pretty standard
- 39:28harness for an agent implementation.
- 39:31Nothing too fancy, but I'm sure they did
- 39:33a ton of work. But I feel like this is a
- 39:35good example for how builder these days
- 39:37should be building products because this
- 39:39is becoming a standard, right? Every AI
- 39:42agent tool should be self-improving and
- 39:44self- evvolving. And depending on the
- 39:46user request, you can either keep it
- 39:48local, you can keep it on the cloud,
- 39:49that's up to you. I feel like the skill
- 39:51and memory accumulation is probably the
- 39:54most valuable thing in this because the
- 39:56user kind of just over time get locked
- 39:58in and it remembers who I am. Which is
- 40:01why like recently I don't switch to
- 40:03other AI anymore. I just use cl code. It
- 40:05has a semantic memory about who I am, my
- 40:07company information, ultimately
- 40:11YouTube videos. But yeah, I think it's a
- 40:13cool framework. You guys should try it
- 40:15out. It's really fun. Cool. I hope this
- 40:17video was helpful. If you have any
- 40:18questions, please leave down a comment.
- 40:20And if you enjoyed it, please give me a
- 40:22like and subscribe. Let me know what
- 40:23else you would like to watch. I will see
- 40:25you next time. Thanks. Hey everyone,
- 40:26this is Sean. So today, let's talk about
- 40:29loop versus graph engineering. Two
- 40:31biggest buzzwords in AI agent space
- 40:33recently. I am probably exactly like
- 40:36most of you guys, which is I don't know
- 40:38how to keep up with all these new
- 40:39concepts, but I still find it very
- 40:41interesting because whenever a new
- 40:43concept becomes viral, usually I'm
- 40:45curious about what triggered it. What I
- 40:47find out is that this guy called Peter
- 40:48Steinberger who's also the author of
- 40:50this little lobster which is open claw
- 40:53said that are we still talking about
- 40:55loops or did we shift to graphs yet? And
- 40:58he posted this on July 18th 2026 which
- 41:01is about 13 days ago and he got 3
- 41:04million views and look at this first
- 41:06comment. It said bro stop I am on
- 41:08vacation. So sometimes you know these
- 41:11important people in the space who
- 41:13mention a keyword and then everybody's
- 41:16going to start talking about it. I
- 41:17remember exactly something like that in
- 41:19context engineering. And for today this
- 41:21video we are going to demystify all of
- 41:23these core AI agent concepts and
- 41:25especially on loop and the graph. We're
- 41:28going to jump in a little bit on the
- 41:29system design real quick. And at the
- 41:31same time, we're also going to walk you
- 41:32through a real coding example uh called
- 41:35Waku- Agent, which is under my GitHub,
- 41:37Shen Sean Chen, and we published this
- 41:39about two three weeks ago, and we got
- 41:41more than 700 stars. If you think this
- 41:43video is helpful for you, please give us
- 41:44a star, give us a like, and that would
- 41:46be very helpful for us. With Waku As
- 41:48agent, you can launch a dashboard like
- 41:49this, and you can just add something
- 41:50like what's up on my Google calendar on
- 41:56Thursday,
- 41:58right? And then it's going to check what
- 42:00kind of graph it's going to use, decides
- 42:02if it needs to retrieve some memories,
- 42:04and then eventually decide, you know,
- 42:06what kind of tools it should be using.
- 42:07Right? Right here, it's listing the
- 42:08events from my Google calendar and
- 42:10eventually send you the reply, and it
- 42:13did a bit of a testing in the vow
- 42:15system. And then, you know, it spit up
- 42:17the gave me the final answer for the
- 42:20results. I'm going to talk a bit more
- 42:21details into this, but without further
- 42:23ado, let's get started on the system
- 42:24design. First thing first, I want to
- 42:26talk about the AI agent engineering
- 42:29ladder that we have come across over the
- 42:31past three years because I think that
- 42:33helps us to kind of understand where we
- 42:36came from and where we're going. And the
- 42:38first one is obviously prompt
- 42:40engineering. And to me, prompt
- 42:41engineering was the very first or early
- 42:44version of how people start to get used
- 42:46to how to use an LLM. I feel the most
- 42:49accurate word to describe this was role
- 42:51playing because at the very beginning,
- 42:52we didn't know how to use LMS. We had
- 42:54chatbt 3.5 and we tell it, hey, you are
- 42:58the best poet like your Shakespeare or
- 43:01Leai, right? Write me a poem about this
- 43:04scenery I'm looking at right now in
- 43:06Switzerland.
- 43:07Then it's going to pretend that it's one
- 43:09of those poets you just mentioned and
- 43:11then talk to you back. Right? That was
- 43:12prompt engineering. You basically draft
- 43:14the prompt so that you control the LM to
- 43:17talk to you in a certain way you want.
- 43:19And then we quickly evolved into what we
- 43:22call context engineering. And that is
- 43:24because when people move from just
- 43:28playing with LM as a consumer product to
- 43:31actual workflows, you realize that you
- 43:33not only need prompt but also you need
- 43:35to feed in data. For example, if you
- 43:38build a customer service chatbot or if
- 43:40you build a sales agent like Automanas,
- 43:43which is my company, you need to really
- 43:44think about how do you construct the
- 43:46context for the agent so that the agent
- 43:49will be able to talk to clients on
- 43:51behalf of the businesses, right? So, you
- 43:54not only will say, hey, you are the best
- 43:57salesperson in the world. You will talk
- 43:59in certain ways. Do not say certain
- 44:01things. You also need to feed in data
- 44:03such as what are some of the what are
- 44:04some of the customer relationship data
- 44:06people already have on their Excel
- 44:08sheets, Google sheets, CRM system, all
- 44:11these kind of things and you need to
- 44:12really put them together and make sure
- 44:14that the agent can accurately be fed
- 44:17with the right information. So we moved
- 44:19from prompt to context very quickly to
- 44:21fix that problem. And then we moved on
- 44:23to skills which I think is sort of
- 44:26teaching the AI what kind of procedure
- 44:29it should follow. Right? It's a
- 44:31procedure following process in AI agent
- 44:34hardness. This is called procedural
- 44:36memory. It's a memory that you tell say
- 44:38you tell a kid, hey, when you walk home,
- 44:42walk on the right hand side if you're in
- 44:44the US or in China, right? Walk on the
- 44:46left hand side if you're in the UK or
- 44:48Japan. That is a procedural memory that
- 44:51you want a person or you want an LLM to
- 44:54remember. And there's no, you know,
- 44:57extra data. It's just fact that it
- 44:59should be doing when a certain situation
- 45:01happens. Okay. So why do we need skills?
- 45:04It's because that if you just p if you
- 45:07if you provide a lot of context to the
- 45:09AI, sometimes it can be a little
- 45:11repetitive. You don't want to always
- 45:14provide you know the same order to AI
- 45:16again and again and again. So having a
- 45:18skill to determine a workflow becomes
- 45:21really really handy. Just like for
- 45:24example if you're coding on claw code
- 45:26you don't want to explain to claude that
- 45:29do not use emoji do not use emoji do not
- 45:31use emoji or I prefer to use emoji use
- 45:34these emojis don't use the brain emoji I
- 45:36hate that right these kind of things you
- 45:38have to repeatedly tell the context then
- 45:40it's becoming less convenient so you can
- 45:42build up skill to make sure the LM is
- 45:45following that procedure then comes with
- 45:47loop what does loop do a loop is
- 45:50basically saying hey Maybe sometimes you
- 45:53have a goal and you want to finish that
- 45:56goal but we don't know the exact skills
- 46:00you should have to finish that goal. We
- 46:02might tell you, hey, there are a bunch
- 46:03of tools you can use, right? That's why
- 46:05Enthropic come up with MCPS, you can
- 46:07call your Google calendar, you can call
- 46:09your Gmails APIs, you can call your
- 46:11GitHub APIs. These tools become handy
- 46:14and then you're telling the LM be like,
- 46:16okay, run a loop, know your goal, which
- 46:19is maybe help me fix this bug that the
- 46:21customer come up with. Run this loop and
- 46:24here are a bunch of tools. Here's a web
- 46:26search tool. Here's a bunch of API
- 46:27calls. Here's a bunch of MCPS. Use them.
- 46:31loop it until at some point you finish
- 46:33the goal and that's the end of the
- 46:35iteration. And then the question comes
- 46:38why do we need graph? What does graph
- 46:41do? Because graph technically is a
- 46:43procedure that is predetermined. Many
- 46:46people are criticizing graph engineering
- 46:47is not new because maybe in 2023 I
- 46:50remember people were already using
- 46:51airflow. people using step functions.
- 46:53People are talking about how do we make
- 46:55sure that a deterministic workflow can
- 46:58be set up properly so that we don't just
- 47:00tell LLM to make all the decisions
- 47:01because sometimes we know exactly how a
- 47:03certain task needs to be done. There's a
- 47:04step there. There's an SOP there. So to
- 47:08me it kind of feels like okay we moved
- 47:10from skills which is strict procedure
- 47:13following to loops which is hey agents
- 47:15go figure it out yourself to eventually
- 47:18we realize that hey we need a mixture of
- 47:20both. That in my opinion is graph. Okay.
- 47:24Technically sometimes you should write
- 47:26the skill first. Okay. How do you cross
- 47:30the road? On what side do you walk in
- 47:32the pavement depending on which country
- 47:33you're in? And do you respond to a
- 47:36client? How do you respond to me when
- 47:38I'm coding with an coding agent? Right?
- 47:40And then you should turn that into a
- 47:41graph. If you see that there's a lot of
- 47:43repetition in the workflows, especially
- 47:45if you're building a workflow for
- 47:46corporate, right? If you're working in
- 47:48uh e-commerce, maybe you see exactly how
- 47:51your customer service should be
- 47:53answering questions related to
- 47:54logistics, to refund, to checking
- 47:56samples, you realize that you can really
- 47:59consolidate this into a graph. When the
- 48:01workflow stops changing, perhaps there's
- 48:03still part of the workflow that you need
- 48:04to use a loop where the loop is
- 48:07basically doing this explorative work
- 48:09out there, right? It's probably doing,
- 48:11you know, a bunch of research for you.
- 48:13Um I think deep research is one of the
- 48:15early examples of a loop engineering
- 48:17workflow here where you just say hey go
- 48:19crazy just go search the internet I want
- 48:21to report these kind of work are you
- 48:24know less standardized or it doesn't
- 48:26have an SOP in it. It's just about I
- 48:29want more information or I have a
- 48:31certain goal use the tools available to
- 48:34you to figure it out for me. is quite
- 48:36different from graph because sometimes
- 48:38we know exactly what tools they should
- 48:40be using to uh finalize a task for me.
- 48:43Okay guys, that was a conceptual
- 48:46walkthrough of this AI agent engineering
- 48:48ladder. Now let's take a look at what a
- 48:50loop and a graph look like. What a loop
- 48:53does is that it discovers what to do
- 48:55next. So maybe this is you and you ask a
- 48:58question to an LLM and the LM is
- 49:00basically saying, "Hey, do I need to use
- 49:02any tools? If yes, choose a tool, run
- 49:05it, and check if it finish the task,
- 49:08come back to the LM, and then loop it
- 49:10again and again and again until at some
- 49:12point LM is like, hey, we don't need the
- 49:14tool anymore. Let's reply to the user.
- 49:17Examples here could be like you are
- 49:19trying to fix a bug on a pull request on
- 49:22GitHub. You basically ask claw code or
- 49:24codeex and be like, fix this bug, tell
- 49:27me when it's fixed. And it's going to go
- 49:28ahead and then use a bunch of tools like
- 49:31web search, github CLI, checking your
- 49:33superbase, checking your AWS, checking
- 49:35your Google cloud, all these kind of
- 49:36stuff and then at the end saying okay
- 49:38we're done. Okay, bug is fixed because
- 49:41of ABCD that is a loop. A graph on the
- 49:44other hand
- 49:46is somewhat similar but uh it's not
- 49:48exactly the same. So you might have a so
- 49:50you might have a standardized process
- 49:52every day and be like oh I want to
- 49:54understand how many people submitted
- 49:55pull requests overnight and who are
- 49:57these guys who submitted some tasks can
- 50:00we take a look at them and uh tell me if
- 50:03uh things are fixed or not and maybe at
- 50:06the same time I want to understand you
- 50:08know what are some of the latest news
- 50:10out there on AI agents and maybe we want
- 50:12to do some web search as well maybe we
- 50:14want to run some you know git commands
- 50:16at the same time we know exactly how it
- 50:18works Okay. And then you probably want
- 50:21this to be done in parallel. So you ask
- 50:24a question and then it's going to check
- 50:25out the GitHub to check out the pull
- 50:27request. It's going to search the
- 50:28website for you to do some explorative
- 50:31analysis. And it's probably also going
- 50:33to check your calendar plus the memories
- 50:35stored locally or on the cloud. And then
- 50:39it's going to synthesize these
- 50:40information and tell us, hey, is there
- 50:42anything else to do? If not, end this
- 50:44and reply. If yes, please explain what
- 50:48kind of things do you still need to do.
- 50:49Okay, you see the difference here. So
- 50:51sometimes you can have some loops here.
- 50:53Okay, maybe having an agent maybe having
- 50:55an agent loop to search the web or have
- 50:58an agent loop to fix the bugs on GitHub
- 51:00is part of this graph. Okay, so you can
- 51:03see graph is basically saying, okay, I
- 51:05know exactly what you should be
- 51:06checking. Maybe they're in parallel,
- 51:08maybe they are happening in sequences.
- 51:10Do it the way I want. And uh once you
- 51:13finish uh synthesize it and tell me the
- 51:15answer. Are you guys still with me?
- 51:17Let's check a real example first. Come
- 51:19back here. Let's come to Waku agent
- 51:22dashboard. The way you set it up, by the
- 51:23way, is come to this website
- 51:25github.com/jennonchan/wacu-agent.
- 51:28You can either click on code and then
- 51:30copy this and then type into your
- 51:33terminal and say get clone and paste
- 51:36that in and hit enter which is this way.
- 51:38And you can just copy this and paste in
- 51:39your terminal. or recently we released a
- 51:42new package in Python called Waku-
- 51:44Agent. All you need to do is copy this,
- 51:47pip install the Wacu agent into your
- 51:48terminal, set up your environmental
- 51:50keys, and then you can launch a
- 51:51dashboard. Copy this,
- 51:54paste this in
- 51:58because I'm already using port 777. So,
- 52:00let's use 778 as an example. Paste that
- 52:04in. You can see this is it. Okay, since
- 52:07I already set it up, I'll come back to
- 52:08this localhost 7777.
- 52:11So, what you're seeing right here is is
- 52:12an entire AI agent harness starting from
- 52:15the gateway, which can be the chat of
- 52:17here, or you can use some other channels
- 52:19such as Discord, Telegram, WhatsApp,
- 52:21stuff like that. And then it's going to
- 52:24check out a retrieval gate to see if we
- 52:26need any procedural memory which is
- 52:28skills as we mentioned or semantic
- 52:30memory or episodic memory which are
- 52:32durable facts or the dated events that
- 52:34happened in your local memories. Okay.
- 52:37And then it's going to run through an
- 52:38agent loop using LM agents and calling
- 52:40the tools and eventually give you the
- 52:42reply during which the LM ops is going
- 52:44to trace the data test it and then
- 52:47release it the new version of the prompt
- 52:49and feed back to the harness. This is
- 52:51the entire harness. Okay. And loop
- 52:53engineering is happening here in this
- 52:54little loop. A graph workflow is
- 52:57basically after this gateway you can
- 53:00predefine some of the process in between
- 53:02here. So we have a new tab called graph.
- 53:04We currently have two graphs here. One
- 53:05is called triage another one is called
- 53:07gather. For triage graph what it does is
- 53:10that it's saying okay start point and
- 53:12then it's going to classify if it
- 53:15requires some agent calls and then at
- 53:16the same time it might just check out my
- 53:18calendar and see what's going on. Okay.
- 53:20And I can just be like what's up today?
- 53:25All right, you can see it triggered this
- 53:27triage first, right? It checked the
- 53:29Google calendar for me and also at the
- 53:31same time it was checking, you know, if
- 53:32we need to have some uh need to do some
- 53:35serious Asian calls here and it
- 53:38eventually decided that okay, it's going
- 53:39to use some tools. So they use the list
- 53:41events read Apple calendar for me. So
- 53:43you just saw a very simple graph kind of
- 53:45call already. Okay. So what it did was
- 53:48that it checked as I mentioned if it
- 53:51needs some serious agent calls and in
- 53:54parallel at the same time it was already
- 53:55checking my calendar because I was
- 53:56asking what's up today and then after
- 53:58that use this loop to run to call the
- 54:00tools like list events read apple
- 54:02calendar and all these kind of stuff and
- 54:04then uh give me the reply. Okay but what
- 54:07if I need to test something else. So
- 54:09here in this local Asian harness what I
- 54:11recently updated is that you can very
- 54:13clearly mention the workflow you built
- 54:16in the graph we have gathered. So you
- 54:18can clearly say slashgather and then say
- 54:22tell me what's up with waku agents and
- 54:24any new PRs any competitive projects.
- 54:27Let's see what happens. You see that it
- 54:29triggered this graph instead. Okay let's
- 54:32come back to the overview. It's showing
- 54:33me that it was trick triggering this
- 54:35graph and it used these tools
- 54:38simultaneously
- 54:40and then it's synthesizing the answers
- 54:42for me. Okay. So, so what it did was
- 54:45that it gave me a morning brief on what
- 54:48kind of PRs are out there and we have
- 54:51released a new package which you can
- 54:53check in this GitHub if you scroll down
- 54:55a little bit. We have released a new
- 54:56package for Asian graphs recently and
- 54:59also there are some PRs to be reviewed.
- 55:01Okay, let's come here. So, if you click
- 55:04into pull request, you can see there are
- 55:06some PRs here for me to be reviewed. And
- 55:08for the web search, you can see that it
- 55:10researched about Harrison Chase publish
- 55:12your harness your memory and your own
- 55:14videos in harness eval ranking. Okay, it
- 55:18was doing the command research for me as
- 55:20well. All right, so you might be
- 55:22wondering how did this work? If we come
- 55:25back to the tab for graph, you can see
- 55:27that we clearly defined two different
- 55:28use cases of graph. One is triage,
- 55:30another one is gather. And they have
- 55:32very specific ways of running the
- 55:35agents. And they both use loops because
- 55:38when it does the web search, it needs to
- 55:39search for the information and come back
- 55:41to me. Right? When it does the check
- 55:43calendar, it's going to check the
- 55:44calendar until it find out what's going
- 55:46on in my calendar and then come back to
- 55:48me. Okay? These things are all kind of
- 55:50intertwined in this agent harness
- 55:52system. So if we come back to the system
- 55:53design, there's something interesting I
- 55:55wanted to show you which is how exactly
- 55:57did the workflow finish from the
- 55:59beginning to the end from a time usage
- 56:02and parallel processing perspective. For
- 56:04the loop, an agent loop, for example,
- 56:06web search is going to decide what it's
- 56:08going to do first, right? It's going to
- 56:10check out my GitHub calling tools one by
- 56:13one, right? Maybe after the first call
- 56:15the agent decided okay we need to do one
- 56:17more call and then it called it again
- 56:19and then it checked the calendar and
- 56:21then it synthesized information for me.
- 56:23But in the graph example which is this
- 56:26which is this workflow for gathering
- 56:29information. It used multiple tools in
- 56:31parallel because we already told it you
- 56:32need these four tools, right? Maybe each
- 56:34one of them spend different amount of
- 56:36time and it did the GitHub first and the
- 56:38web search did the most amount of time.
- 56:40Calendar check and memory retrieval did
- 56:42the least amount of time because it's an
- 56:44instant or maybe because it did not need
- 56:46any memory retrieval at all. And then
- 56:48it's going to synthesize the information
- 56:50because it's got a lot of context at the
- 56:52same time. So you cannot say so you
- 56:54cannot say graph engineering is
- 56:56replacing loop engineering because
- 56:58they're coexisting and you cannot say
- 57:00loop engineer graph engine is better
- 57:01than the other one because we need them
- 57:03in different use cases. A loop is
- 57:05something you need when the model
- 57:06decides what to call one step at a time.
- 57:09A graph is something like when you know
- 57:11the shape and you just want them to move
- 57:14together. Okay, it really depends on the
- 57:16situations you're building for. In your
- 57:18local files, we have a WKU folder and
- 57:21then we have a graph folder and within
- 57:24the graph folder, we have defined the
- 57:27graph engines which is literally a graph
- 57:31and you're going to add nodes which in
- 57:33our case could be web search tools,
- 57:35could be agent calls, could be MCPS,
- 57:37anything right and then you're also
- 57:39defining edges which is okay do we where
- 57:41do we go after using one tool right
- 57:44after the web search do we go
- 57:45synthesizing or do we go somewhere else
- 57:47Right graph we also defined nodes which
- 57:51is basically saying what kind of things
- 57:53can be a node right to calls LLM calls
- 57:57agent calls router these kind of things
- 58:00and then under this graph folder we also
- 58:02have another folder called workflows you
- 58:04can see we have triage here right
- 58:06remember triage was the chart we showed
- 58:08you which is it's going to classify if
- 58:10it needs to use some complex models to
- 58:11do agent calls or is it doing some
- 58:14calendar checking at the same time these
- 58:16two things are happening in parallel at
- 58:18the same time. And the close is actually
- 58:20very simple. You define these
- 58:22functionalities to make sure that the
- 58:24triage graph is working properly with
- 58:26this predefined workflows, right? You
- 58:29add node when you need to add more
- 58:31tools. You add edges by defining the
- 58:34initial problem from start can go to
- 58:36either classify or check calendar at the
- 58:39same time. If we check out the gather
- 58:42tool as well, remember it's a similar
- 58:44thing as our preview chart. You can scan
- 58:46the GitHub website calendar memory.
- 58:50Okay, GitHub website calendar memory.
- 58:53These are all the nodes and after that
- 58:55you synthesize it and then you decide if
- 58:57we should return the answer. Obviously
- 58:59we're doing the same thing. We're
- 59:00building the graph here, right? We're
- 59:02adding these nodes and we are adding
- 59:05those edges too. I hope this is easy to
- 59:07understand and you can feel free to add
- 59:11more.py files here to make this graph
- 59:13workflows even larger. And in that case,
- 59:16I would love to work with you and then
- 59:17you can feel free to contribute to this
- 59:19repo and be one of our contributors and
- 59:22submit your own workflows because once
- 59:24you submit your own workflows and it's
- 59:26approved, it will be available here. If
- 59:28I type in slashgraphs,
- 59:31it's going to tell me that it has
- 59:34gather, it has triage and triage is the
- 59:36router. It's basically going to read the
- 59:38local files of these workflows from
- 59:40graphs and people can use it. This will
- 59:43be very cool and then we'll probably
- 59:44become a community if you guys are
- 59:46interested and feel free to be our
- 59:47contributors and also if you're
- 59:49interested in talking more in depth of
- 59:51these concepts with us if you feel you
- 59:53want to discuss with me about your
- 59:54workflows and any questions you have
- 59:56about system design about these agent
- 59:57hardness and all these kind of stuff you
- 59:59can feel free to come to my personal
- 1:00:00website shantchan.io O and then click on
- 1:00:03join over here to join our community.
- 1:00:05I'm just getting started to do this
- 1:00:07because I don't have enough time to
- 1:00:09answer all of your questions. Feel like
- 1:00:10the most efficient way is that I can
- 1:00:12host some live sessions with you uh
- 1:00:13twice a month so that I'll be able to
- 1:00:15answer most of your questions and we can
- 1:00:17prepare some you know build sessions
- 1:00:19real time showing you my real setup and
- 1:00:21all this kind of stuff. If you join this
- 1:00:22community, you will be assigned to a
- 1:00:24private discord and I will share more
- 1:00:27details there including the original
- 1:00:29files of all of these system design
- 1:00:31charts that I built in the past uh in
- 1:00:34real code so that you can sort of open
- 1:00:36it in your own scala draw website to
- 1:00:38learn about it. So last but not least, I
- 1:00:40think this is an important question we
- 1:00:41should ask ourselves which is isn't this
- 1:00:44just a deterministic workflow from 2023?
- 1:00:46What is new now is that some of the
- 1:00:48nodes that we're using right now are not
- 1:00:49deterministic. could be a LM call and at
- 1:00:52the same time sometimes the model can
- 1:00:54pick the edge right the routing is the
- 1:00:56way that you know you're letting the LM
- 1:00:59as a judge do I pick a simple model to
- 1:01:02answer the questions or do I use a more
- 1:01:05complex model to answer this question
- 1:01:07and also you need some guards which a
- 1:01:09previous directional graph scheduler
- 1:01:12never did these are buzzwords as I
- 1:01:14mentioned but buzzwords are viral for a
- 1:01:16reason sometimes because of famous
- 1:01:18people sometimes because it's actually
- 1:01:19useful So I hope this kind of videos is
- 1:01:22helpful for you and again if you have
- 1:01:24any questions feel free to ask me and I
- 1:01:27would love to answer your questions live
- 1:01:28in our community just come to my
- 1:01:30personal website shanten.io and then
- 1:01:32there are a bunch of sources here
- 1:01:34looking forward to and if you love this
- 1:01:36project please give us a star on GitHub
- 1:01:37repo and I would love to work with you.
- 1:01:40Thank you so much for your attention.
- 1:01:42Appreciate it.
About this transcript
This page contains the full transcript of 8 сентября 2026 г. by Andrei, generated from the public captions YouTube serves with the video. The transcript has 12,711 words across 1,785 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.