همهچیز دربارهٔ AI Agent — از مدل زبانی تا ایجنت خودمختار — Transcript
Full transcript
- 0:01Everyone is talking about AI agents
- 0:02these days.
- 0:15But the reality is that most of us are
- 0:17still using AI the same way we did a
- 0:19year ago. For example, we type a
- 0:23question into ChatGPT or Claude, get an
- 0:25answer, copy it, and use it elsewhere,
- 0:28but we can use AI for so much more now.
- 0:32By using agents, we can do much, much
- 0:34more interesting things. See, a chatbot
- 0:38only lives inside that chat window. You
- 0:42ask it a question, and it gives you an
- 0:44answer. It has no access to other tools
- 0:46. In contrast, an agent has that same
- 0:50brain, but imagine it also has hands
- 0:52and feet. It has access to a set of
- 0:57tools, access to your files, access to
- 0:59the web so it can search, it has some
- 1:01memory, specific goals, and infinite
- 1:03tools you can connect to it. In this
- 1:12video, I want to show step-by-step what
- 1:14an agent is capable of, what its
- 1:15internal anatomy is, what role each
- 1:17part of an agent plays in its
- 1:19performance, and knowing this will help
- 1:21you understand where you can use an
- 1:22agent in your work, what you need to
- 1:24build one, and this is the most
- 1:26important step for optimally
- 1:27integrating AI into your work and
- 1:29projects. Well, look, the foundation of
- 1:37all these tools is one thing: a Large
- 1:39Language Model, or an LLM, which acts
- 1:41as the brain of this system; like
- 1:44ChatGPT, Gemini, Claude, and all of
- 1:46those. They are applications built on
- 1:50top of a language model, and they work
- 1:51simply. You give it an input, and the
- 1:55model gives you an output based on the
- 1:57data it was trained on. For example,
- 2:01you ask it to write an email on a topic
- 2:03you choose; your prompt is the input,
- 2:05and the email text you receive, which
- 2:07is usually more polite than what we
- 2:08write ourselves, is the output of this
- 2:10language model. Now, ask that same
- 2:14model when my next meeting is, and it
- 2:16cannot answer. Why? Because it doesn't
- 2:18have access to your data. It doesn't
- 2:21have access to your calendar, and this
- 2:23actually highlights two important
- 2:24characteristics of these language
- 2:26models. Two features that are very
- 2:29important for understanding agents. One
- 2:32is that these models are indeed trained
- 2:34on a massive amount of data; from the
- 2:36whole world, from all sources. But they
- 2:39know nothing about your private
- 2:41information, your calendar, your emails
- 2:43, or your internal company documents.
- 2:46It’s like a brain that knows all of
- 2:48human knowledge, but it doesn't know
- 2:50you specifically, nor does it know the
- 2:51details of your life. Secondly, these
- 2:56language models are passive; meaning
- 2:58they sit and wait for you to give them
- 3:00a prompt, ask a question, so they can
- 3:02respond and provide the text answer to
- 3:04you. It doesn't start anything on its
- 3:07own, and because it lacks access to
- 3:09tools, it can't perform any actual
- 3:11tasks. As a very simple example,
- 3:14imagine a highly professional chef
- 3:16sitting in an empty room. Now imagine
- 3:21that same chef running the kitchen in
- 3:23the best restaurant in town. Why?
- 3:26Because they have access to various
- 3:28tools like the stove, fresh ingredients
- 3:30in the fridge, and professional cooking
- 3:32utensils, allowing them to prepare the
- 3:34best dishes. But you can only go to the
- 3:36chef sitting in the empty room and ask
- 3:38questions like, "What should I do for
- 3:40this dish?""What do you think about
- 3:42this?" Consequently, this is the
- 3:45difference between an AI agent and a
- 3:47regular chatbot like ChatGPT or Claude.
- 3:50Of course, Claude and ChatGPT have
- 3:52advanced significantly now and their
- 3:54agentic capabilities have increased,
- 3:56but we are using their chat mode as an
- 3:58example. Before we get to agents
- 4:00themselves, let's have a general
- 4:02classification together. The use of AI
- 4:07can be categorized into 4 levels, and
- 4:09most people, without knowing higher
- 4:11levels exist, remain at the first level
- 4:13, which is using chatbots. So, level
- 4:17one is chat. Meaning you ask a question
- 4:20and get an answer. This is very useful
- 4:23and has already transformed our lives
- 4:25significantly. But you have to do the
- 4:27rest of the work yourself. You have to
- 4:29make the decisions. When to ask a
- 4:31question? What to ask, and then you
- 4:33have to copy the text, take it
- 4:34somewhere else, and do whatever you
- 4:36need to do; you have to handle the rest
- 4:38yourself. Level two is AI tools.
- 4:41Meaning tools that can perform specific
- 4:44tasks, for example. For instance, you
- 4:47see an AI tool dedicated to generating
- 4:49images. One tool is dedicated to
- 4:52creating slides. One tool is dedicated
- 4:55to analyzing tax documents, and
- 4:57ultimately, they deliver a final
- 4:59product. But you are still the one
- 5:02behind the wheel. Meaning you decide
- 5:04yourself, "I need to go use this tool
- 5:06right now.""Delegate this task to it.""
- 5:10Get this output from it." What is level
- 5:12three? AI workflows; this is where
- 5:14tasks are chained together. Meaning you
- 5:20define a series of specific tasks you
- 5:22want performed in a sequence—that is,
- 5:24"do this, then do that, then do this"
- 5:26—within a pipeline, and from then on,
- 5:28it automatically executes them one
- 5:30after another. For example, assume that
- 5:36every morning you take links to several
- 5:38news items from LinkedIn, ask a model
- 5:40to summarize them, and then ask other
- 5:42models to write a LinkedIn post based
- 5:44on that summary. This is an example of
- 5:50a workflow where you perform a series
- 5:52of different tasks in a chain, each of
- 5:53which might even require AI. Meaning,
- 5:59you can build such a workflow using
- 6:01tools like n8n or Make and set it to
- 6:03run for you every day at 8 AM. The
- 6:08important feature of a workflow is that
- 6:09it can only follow a path that a human
- 6:11has predefined for it. For instance, if
- 6:16you ask a calendar workflow about the
- 6:18weather, well, it cannot answer.
- 6:21Because the path you defined only goes
- 6:22through the calendar. It has no access
- 6:25to anything else. In short, when a
- 6:27human has decided what the workflow
- 6:29path should be, it is no longer an
- 6:31agent. It is a workflow. Now, before we
- 6:35reach the fourth level, or agents,
- 6:37let's talk about the concept of RAG.
- 6:40RAG is an abbreviation for
- 6:42Retrieval-Augmented Generation. What is
- 6:44its general concept? Since we are going
- 6:46through this video in a summarized
- 6:48manner. The general concept is that you
- 6:52provide a database or an information
- 6:54source to the AI you are using. This
- 7:02means you can provide specific
- 7:04information—in the form of PDFs,
- 7:06databases, tables, images, or anything
- 7:08you can think of—in its own specific
- 7:10format to the language model and ask it
- 7:12to refer to that database before
- 7:14answering your question, and provide an
- 7:16answer based on that information. Now
- 7:23we reach level 4, which is agents. The
- 7:28only major change that needs to happen
- 7:30for a workflow to become an agent is
- 7:32that the human decision-maker must be
- 7:33replaced by a language model. With an
- 7:38agent, you no longer give it a
- 7:40step-by-step workflow path. You entrust
- 7:43it with a goal. Meaning, you don't say
- 7:45do this first and then do that. You say
- 7:47, "I want this result.""These are the
- 7:49tools at your disposal." The agent
- 7:52itself then finds the path to reach
- 7:53that goal. Structurally speaking, an
- 7:57agent is just a language model
- 7:58connected to four things. First, the
- 8:01language model itself, which is its
- 8:03reasoning engine or its brain. Second,
- 8:07the tools we provide, such as the
- 8:09terminal, browser, file systems it can
- 8:11access, and APIs it has access to—
- 8:13these are like the system's hands with
- 8:16which it can perform tasks. Third,
- 8:19memory. This is very important; it's
- 8:23like information the agent reads at the
- 8:25start of each session so it doesn't
- 8:27start from scratch—a kind of memory
- 8:29of previous tasks performed and
- 8:31previous conversations had, so it can
- 8:33be assigned longer tasks. And fourth:
- 8:38the goal; which means specifying what
- 8:40result you want to achieve, with a
- 8:42clear definition of that result or
- 8:43destination you provide it. One
- 8:48important point that brings agents to
- 8:50life is, in a way, the loop. See, every
- 8:54agent on every platform works with a
- 8:56three-step loop. Meaning the agent
- 9:01observes the current state: what files
- 9:03exist, what image a webpage is showing,
- 9:05or what result the last command
- 9:07returned. Then "thought"; meaning based
- 9:13on what you just said and the current
- 9:15conditions, it thinks and decides what
- 9:17the right next step is. And "action";
- 9:24changing something using one of the
- 9:26tools at its disposal; should it write
- 9:28a file, perform a web search, or
- 9:30execute a code command? And this cycle
- 9:36repeats until it reaches the result and
- 9:37the goal we initially set. If you see
- 9:43this term "ReAct" written with a
- 9:45capital A and a lowercase a, that is
- 9:47the "React" programming language; but
- 9:49if you see it with a capital A, it is a
- 9:51combination of the two words "Reason"
- 9:53and "Act," because this is the
- 9:54combination that agents perform. They
- 9:58reason, they take action. They reason,
- 10:01they take action, and this is the most
- 10:03common framework for agents. Well, so
- 10:06far we have seen the skeleton of an
- 10:08agent and what it is composed of. Let's
- 10:12dive a little deeper into some of its
- 10:14parts. The memory section is an
- 10:17important part of agents. We said that
- 10:20a language model doesn't know you
- 10:21specifically. But agents usually have
- 10:24four types of memory. When I speak of a
- 10:27language model, I am not specifically
- 10:29talking about ChatGPT or Claude;
- 10:31because ChatGPT and Claude are not just
- 10:33raw language models. They also have a
- 10:36memory of you; if you notice, they
- 10:38gradually learn and understand more
- 10:40things about you. When I talk about a
- 10:43language model, the raw language model
- 10:46is ChatGPT. If you have used ChatGPT,
- 10:50Claude, and similar tools, these also
- 10:52have some agentic capabilities added to
- 10:54them; especially Claude, if you have
- 10:56used its software, it has a very, very
- 10:59strong agentic framework. Now let's go
- 11:03back and talk about the types of memory
- 11:05so it becomes clearer for you. The
- 11:10first memory is working memory; every
- 11:12time you talk to the agent, your
- 11:14question, the history of this
- 11:15particular conversation, all the Q&A
- 11:17between you and the chatbot, and the
- 11:19instructions given to the system
- 11:21beforehand are placed in a temporary
- 11:23package; similar to computer RAM. This
- 11:28is a volatile memory. Meaning when the
- 11:30conversation ends, these chats are
- 11:32cleared. Ordinary chatbots only have
- 11:34this. For example, if you go to a site
- 11:37and talk to its chatbot support, close
- 11:39the page, and come back, everything
- 11:41starts from scratch because that memory
- 11:43only held the memory of that particular
- 11:45chat. The second type is procedural
- 11:48memory. This memory actually tells the
- 11:51agent how to behave. For example, a set
- 11:54of instructions, rules, or workflows is
- 11:57given to it, usually as a text file
- 11:59with an extension. It is saved as MD,
- 12:03which is called a skill. You might hear
- 12:06this term a lot these days. This skill
- 12:09is actually a text file in the format
- 12:11of that same Markdown or. MD file,
- 12:14teaching the agent how to perform a
- 12:16specific task. The third type of memory
- 12:19is semantic memory. A series of stable
- 12:22facts that we want the agent to always
- 12:24know. For example, who are you? What is
- 12:28your product? What questions do
- 12:29customers usually ask? Exactly the
- 12:33things that a language model doesn't
- 12:35know about you, but you want your agent
- 12:36to know. And the fourth type of memory
- 12:43is episodic; meaning a time-stamped
- 12:45history of information; meaning that
- 12:47what was said at any given time and
- 12:49what happened are recorded, and the
- 12:50agent is aware of them, like a logbook
- 12:52where events are recorded by time. For
- 12:59instance, suppose you have a problem
- 13:00and you tell your agent. Your agent
- 13:03solves it and records it somewhere. It
- 13:08says, for example, on such and such a
- 13:10day, this issue was resolved, and later
- 13:12if you refer to it, since your agent
- 13:14has this memory, it says this problem
- 13:15occurred on such a date and was solved,
- 13:17and this can be implemented for
- 13:19everything. For all types of dated
- 13:24event logs; for example, if it is
- 13:26inventory management, or an educational
- 13:28topic, or a legal matter, and well, you
- 13:30can imagine how many places it is used.
- 13:37Before moving to the next section, let
- 13:40me mention one point again: if you need
- 13:42consultation regarding integrating AI
- 13:44into your workflow, your projects,
- 13:45building agents, or new AI projects,
- 13:47the way to contact me for consultation
- 13:49is in the description of this video.
- 13:56You can act through my website and
- 13:58schedule a consultation session. Now,
- 14:01these memories are stored in a series
- 14:03of databases. Whenever the agent needs
- 14:08to find one of these pieces of
- 14:10information, it refers to the part that
- 14:11it knows is programmed for that type of
- 14:13information. For example, among the
- 14:191,000 conversations you've had with the
- 14:21agent, it can find 20 specific
- 14:23conversations about a certain topic,
- 14:25and it matches the meaning of that text
- 14:27with what it's looking for; it doesn't
- 14:29necessarily have to search and find the
- 14:31exact same words. Now, an important
- 14:37point is that from a certain point on,
- 14:39the text of the conversations between
- 14:41the person and the agent grows; meaning
- 14:43it might have 1,000 lines of
- 14:45conversation, and as this volume gets
- 14:47larger and larger, it becomes expensive
- 14:49and the agent's accuracy decreases. As
- 14:57a result, what they do after a certain
- 14:59point—for example, once a month
- 15:01passes—is summarize all the previous
- 15:03month's conversations (it's called "
- 15:05distilling," I think you could also
- 15:07call it "refining"), extract the
- 15:09essence, derive a few summary sentences
- 15:11, pull out the key points, and store
- 15:13those specific sentences in their
- 15:15semantic memory. This is what they call
- 15:20memory consolidation. They consolidate
- 15:23certain things into memory, which in
- 15:25turn makes the language model run more
- 15:27cost-effectively. Now, these agents
- 15:30have a kind of memory file. You might
- 15:34have seen this if you have worked with
- 15:35agents. For some tools, it is called
- 15:38cloud.md. For some tools, it is called.
- 15:42md. There are actually a set of rules
- 15:47in this file that the agent reads every
- 15:49time it wants to start talking or
- 15:50working, so it is aware of them before
- 15:52it begins. For a very simple example,
- 15:57suppose you tell an agent to, say,
- 15:59build a website for you. Every time it
- 16:02builds a website, it puts an emoji next
- 16:04to every sentence, and you have to keep
- 16:06deleting them. But if you ask it to
- 16:11record this rule in its memory file—
- 16:13that in the text you put on the site or
- 16:15in the comments, never use emojis—it
- 16:17will never repeat that mistake again,
- 16:19or even more professionally, if you
- 16:21manually know where that cloud.md file
- 16:23is. Where is the md file? You can add a
- 16:29sentence to it telling it not to do
- 16:31that anymore or to do it a certain way,
- 16:33and the agent then learns to structure
- 16:34its work based on the rules you have
- 16:36defined. Now, let's talk a bit about
- 16:41some terms that have recently become
- 16:43very popular regarding agents or that
- 16:45we hear quite often. One of these terms
- 16:49is "harness." Literally, harness means
- 16:53a bridle like a horse's harness, and
- 16:55here it is used as a model harness. If
- 17:01you imagine a language model as a
- 17:03powerful horse, it might run anywhere
- 17:04without control, but with a harness,
- 17:06you can control it to do exactly what
- 17:08you want. Now, why is control necessary
- 17:12? Look, at their core, language models
- 17:16are fundamentally predicting the
- 17:17probability of the next word. This is
- 17:21how they write a long text. Wherever
- 17:23there is probability, there is also the
- 17:24possibility of error. In fact, where
- 17:28you are just chatting, it doesn't
- 17:29matter much, but where the task you are
- 17:31performing is important—you might
- 17:33want your agent to modify files, make
- 17:35certain decisions, or perform specific
- 17:37actions—the probability of this error
- 17:39must be minimized as much as possible.
- 17:44Harnessing is essentially all the
- 17:46guardrails we place around a model to
- 17:47ensure it behaves exactly the way we
- 17:49want. Harnessing is a broad topic;
- 17:53I’ll probably make a separate video
- 17:55for it. But for now, just know that
- 18:01harnessing can include specifying what
- 18:02the model should or shouldn't do in its
- 18:04system prompt, exercising caution when
- 18:06defining tools, the instructions you
- 18:08provide, and managing the loops that
- 18:10repeat until the desired result is
- 18:12achieved. Keep this general idea in
- 18:18mind; I’ll try to discuss it more
- 18:20later. If you're interested in loops,
- 18:25that’s also an important term
- 18:26currently discussed regarding agents.
- 18:33Engineering these loops and setting
- 18:35stop guardrails are also key topics in
- 18:36building and using agents. I’ve
- 18:42mentioned before that an agent rotates
- 18:43through an observation, thought, and
- 18:45action loop; but the important question
- 18:47is, when should this loop stop? The
- 18:50simple answer is when it reaches its
- 18:52goal. But how does the agent know it
- 18:54has reached its goal? It might call a
- 18:58tool 20 times, fail to reach the goal,
- 19:00and then be stuck calling that tool
- 19:01infinitely. Consequently, it needs to
- 19:07know at some point how "good" is enough
- 19:09—that is, what result is sufficient,
- 19:11or if it can't reach a conclusion, we
- 19:13need to define for it when to stop the
- 19:15loop. For example, you tell your agent:
- 19:21"See what customers are complaining
- 19:23about; if anyone has a refund request
- 19:25that hasn't been processed, follow up
- 19:27on it." The agent first pulls data from
- 19:32the customer management system; for
- 19:34instance, suppose there were 30
- 19:35complaints in the last 2 months, 12
- 19:37refunds were processed, and 8 remain.
- 19:41Then it thinks to itself that it was
- 19:43asked to follow up on the remaining
- 19:45cases; should I follow up with these 8
- 19:46people now? Should I schedule a meeting
- 19:49with them? Should I send them an email?
- 19:52Should I start the refund process
- 19:54myself? And well, this is an open-ended
- 19:56scenario. The important point is that
- 19:59for every process, you must anticipate
- 20:01and define an endpoint. If you want to
- 20:07prevent errors, a good agent is one
- 20:08that, when you ask it for a task, asks
- 20:10you during the planning phase: "What is
- 20:12the end goal in your opinion?" Should I
- 20:17process the refund myself, or just
- 20:19deliver the list? Here, you are the one
- 20:22defining the loop's termination
- 20:24condition. This is the stage agents are
- 20:26currently in. Certainly, within a year,
- 20:30they will be able to make these
- 20:31decisions much more easily, and in a
- 20:33way, they will give you choices based
- 20:35on your preferences. I'll give you
- 20:38another important example of loop
- 20:39engineering. Look, consider a situation
- 20:44where an agent might ask you a question
- 20:46to plan or request permission to access
- 20:48something, and you've left the agent to
- 20:50do the job while you've gone off; this
- 20:52might have happened to you; then, for
- 20:54instance, you return after half an hour
- 20:56, happy to see the result, only to find
- 20:58the agent has just asked, "Can I have
- 20:59access to this file or not?" And it's
- 21:04still waiting for file access
- 21:05permission; essentially, half an hour
- 21:07or several hours of your time have been
- 21:09wasted. How, how can you prevent this?
- 21:14In the same harness or file, define a
- 21:16rule for it: whenever you are waiting
- 21:18for my permission for something, send
- 21:20me a notification on Telegram. You know
- 21:26, these are things that can be defined
- 21:27to optimize the agent's running process
- 21:29as much as possible. Well, a question
- 21:33that arises is, we build an agent. It
- 21:36has memory, tools, and its loop is also
- 21:38refined. How do we know this agent
- 21:41works well? Here we reach the topic of
- 21:44LLMOps, similar to DevOps. The "Ops"
- 21:48part stands for operations, meaning
- 21:50language model operations. In fact,
- 21:56evaluation of language models starts
- 21:58with tracking; meaning everything from
- 21:59the moment the user asks a question to
- 22:01the answers the agent gives, the tools
- 22:03it calls, and the thoughts it has, must
- 22:05be recorded. How many times each tool
- 22:11was called, how long each step took,
- 22:13how many tokens were consumed—these
- 22:15must be recorded. The second step is
- 22:18evaluation. Evaluating to see if this
- 22:21execution was good, if it was done
- 22:22correctly. Interestingly, you can even
- 22:27use another language model to evaluate
- 22:29the quality of an agent's work. Meaning
- 22:34you use another language model to
- 22:36evaluate whether this agent performed
- 22:38its task well relative to what was
- 22:39requested or not. Or you yourself can
- 22:43objectively decide whether it did a
- 22:45good job or not. For example, if you
- 22:48had asked it to follow up on customer
- 22:50work, did it do it the way you wanted
- 22:51or not? If it was supposed to send an
- 22:53email, did it send it or not? The third
- 22:55step is diagnosis and correction.
- 22:57Meaning, for example, you detect that
- 22:59this agent took too long to finish its
- 23:01task. Consequently, you must diagnose
- 23:04the reason for this excessive duration
- 23:06there. Has the working memory become
- 23:08too large? Maybe, for a simple question
- 23:11, it's searching the entire system
- 23:13memory. Perhaps it only needs access to
- 23:16a small portion of the memory. That's
- 23:18where the memory system needs to be
- 23:20engineered and fixed. Or maybe a
- 23:22question doesn't require memory
- 23:23retrieval at all. For example, what is
- 23:26the capital of France? The model no
- 23:28longer needs to check its memory; the
- 23:30LLM itself already holds that
- 23:31information. Or perhaps the prompt or
- 23:35instructions you give your agent, for
- 23:37example, saying: "Don't overcomplicate
- 23:39it, just use a simple approach," or in
- 23:41cases where you need more complex
- 23:42analysis, you should tell it to "think
- 23:44more deeply." These are the parts where
- 23:48, after reviewing, you identify a
- 23:50series of flaws and correct them. So,
- 23:53the complete LLM-Ops cycle is tracking,
- 23:56evaluation, identification, and
- 23:58correction. If your agent is built well
- 24:02, it's a system that can even improve
- 24:04itself over time; this is where strong,
- 24:06real agents are separated from weak
- 24:08ones. Agents that can improve
- 24:13themselves over time. Now, this is
- 24:17actually one of the most practical
- 24:18parts of the video that doesn't get
- 24:20much attention. The way you write
- 24:25prompts for an agent should differ from
- 24:26the way you write prompts for a chatbot
- 24:28. When you write a prompt for a chatbot
- 24:33, you are essentially describing what
- 24:34you want from the chatbot. But when you
- 24:39want to give a prompt to your agent,
- 24:41it's as if you are signing a contract
- 24:42with it. This means you must provide
- 24:48precise instructions based on which it
- 24:50is obligated to perform the work and
- 24:51deliver it to you. This is very
- 24:55important. The more precise the
- 24:59instructions and the more clearly you
- 25:01define potential cases—telling it "in
- 25:03this scenario do this, and in that
- 25:04scenario do that"—the more smoothly
- 25:06your agent will operate. For instance,
- 25:12let me give an example: you tell a
- 25:14chatbot to build a website page for my
- 25:16new product. Well, you’ve said
- 25:20something general, and it's also very
- 25:21vague. It can fill in the blanks you
- 25:25haven't fully specified however it
- 25:26wants and get quite creative; it
- 25:28essentially just gives you the code.
- 25:33But give the same text to an agent, and
- 25:36well, the agent has tools, loops,
- 25:38access to various resources, and
- 25:39authority; it starts building a
- 25:41framework it deems appropriate, which
- 25:43might take half an hour. Now, depending
- 25:49on the task you ask of it, one job
- 25:51might take 5 minutes, another might
- 25:52take an hour or two of your time, and
- 25:54in the end, it might turn out to be
- 25:56something you never wanted at all; it
- 25:58depends on how specifically you know
- 26:00what you want and it will cost you. As
- 26:04a result, it is very important that
- 26:06your prompt is like a contract; specify
- 26:08exactly what things it can use. What
- 26:12systems should it use? What frameworks
- 26:15should it not use? Ultimately, the more
- 26:19precisely you write this contract, the
- 26:21more useful an agent with fewer errors
- 26:23you can have. The issue is clearly not
- 26:27just about longer prompts. The prompt
- 26:30should be structured. A structured
- 26:32prompt has four parts. First, its goal
- 26:36must be completely clear; not just what
- 26:38task it should do, but what "finished"
- 26:40looks like exactly from your
- 26:42perspective. What do you want the final
- 26:45output to look like? The second point:
- 26:47constraints, what it is not allowed to
- 26:50do; for example, "do not install new
- 26:51packages without my permission," or "do
- 26:54not make changes to the database
- 26:55without my approval." Whenever
- 27:01something wasn't on your list of
- 27:03constraints and you later see your
- 27:04agent performed that action, you can
- 27:06add those things—which you think of
- 27:08later—to the agent's constraint list;
- 27:10in a way, much of this comes with
- 27:12experience, and it can vary across
- 27:13different tasks. And another point:
- 27:18what is the exact format or structure
- 27:20of the output? For instance, the output
- 27:22you want delivered to you, do you want
- 27:24it in PDF format? Do you want it to be
- 27:26an HTML file? Do you want it saved in a
- 27:28database? You must specify these things
- 27:32completely and think about the points
- 27:34where it might fail; meaning, where
- 27:35might it get stuck? What should it do
- 27:38when it gets stuck? Should it stop
- 27:40completely when it gets stuck? Should
- 27:43it ask you for input? Should it send
- 27:45you a notification on Telegram, like in
- 27:47the previous example? These are things
- 27:50you can define, and it is very
- 27:51necessary to define them for agents. So
- 27:54, knowing this, where should we start
- 27:56building an agent? Look, depending on
- 27:59who you are and what you want from an
- 28:00agent, there are different paths. I
- 28:04don't intend to teach it in this video,
- 28:05but I will mention the different modes.
- 28:09Write in the comments which one is more
- 28:10interesting to you so I can teach it in
- 28:12the next video. Whichever gets more
- 28:14votes, I will definitely make a
- 28:16tutorial video for it. If you are
- 28:20someone who wants an all-purpose agent
- 28:22on your own desktop computer, tools
- 28:24like Anthropic's Claude Code or
- 28:25OpenAI's Codex are the best options for
- 28:27you. Meaning they are agents themselves
- 28:32. They have access to many tools
- 28:34themselves. You can also add many tools
- 28:37to them later yourself, and they can
- 28:38perform many tasks for you. From
- 28:41organizing files and extracting PDFs to
- 28:43building things and creating apps, they
- 28:46can do many of these tasks. But if you
- 28:51are into building workflows, for
- 28:52example, you can use tools like n8n or
- 28:54Make. I have taught n8n in two of my
- 29:00videos. In n8n, you can build more than
- 29:03just workflows. You can build agents.
- 29:05If you haven't seen those videos of
- 29:07mine, definitely go watch them. I will
- 29:08put the link in the description. If you
- 29:12want to build a personal assistant
- 29:14without any coding that, for example,
- 29:16checks your emails or replies on
- 29:17WhatsApp. Platforms like Flowise and
- 29:23Automa. Can be useful for you. The
- 29:28Automa platform. I also talked about it
- 29:31in my previous video. If you're
- 29:33interested in Basecamp 4, let me know
- 29:34and I'll make a video about it. But if
- 29:38you're a developer and want to dive
- 29:40deep into the technical side to build
- 29:42exactly what you need, and of course
- 29:44build much more powerful things, you
- 29:46can create your own agents using
- 29:47frameworks like LangChain or Pydantic
- 29:49AI. Define your own local models. Use
- 29:54local models. Assemble the components
- 29:57yourself and, in fact, keep control
- 29:59over all the details. Especially
- 30:03Pydantic AI, which in my opinion is
- 30:05very, very strong for agentic systems
- 30:07right now. Where LangChain and
- 30:10LangGraph show certain limitations,
- 30:12Pydantic AI is quickly resolving them
- 30:14and moving forward. Now, neither of
- 30:17these methods is better than the other.
- 30:20It just depends on what you need from
- 30:22building an agent, so choose based on
- 30:24that. Which one of these methods to use
- 30:27. Let's do a short recap. Three
- 30:31important things we learned in this
- 30:32video. First, about agent architecture:
- 30:38an agent is essentially a language
- 30:40model that, besides tools, also has
- 30:42memory and a goal, which are connected
- 30:44by a loop of observation, thought, and
- 30:46action. Now, the next time someone says
- 30:52"AI agent," you'll know exactly what
- 30:54they're talking about. Second, the
- 30:57discussion on memory was important;
- 30:59memory can be a simple text file that
- 31:01you place alongside your agent. It
- 31:06could be data that is continuously
- 31:07recorded in your database, which your
- 31:09agent has access to. And third, the
- 31:17prompt contract, which is very
- 31:18important because it defines the
- 31:20agent's goal, its constraints, the
- 31:22output format, and what it should do in
- 31:24case of failure or errors. These were
- 31:30the three important things about agents
- 31:32. I hope you found it interesting. I
- 31:36would like to continue this series on
- 31:37agents. To teach more about it. And the
- 31:41various tools available for building
- 31:43agents. Be sure to write in the
- 31:46comments exactly what interests you for
- 31:47future videos. If you need my
- 31:52consultation regarding your projects,
- 31:54building agents, or creating automation
- 31:56workflows for your company or work, the
- 31:57link to contact me is in the
- 31:59description. You can definitely use
- 32:04that link through my site to get in
- 32:05touch with me, and as always, thank you
- 32:07for watching. For now, until the next
- 32:11video, goodbye.
About this transcript
This page contains the full transcript of همهچیز دربارهٔ AI Agent — از مدل زبانی تا ایجنت خودمختار by Maryam Sadeghi, generated from the public captions YouTube serves with the video. The transcript has 4,904 words across 738 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.