YouTube2Text

همه‌چیز دربارهٔ AI Agent — از مدل زبانی تا ایجنت خودمختار — Transcript

by Maryam Sadeghi · 4,904 words · 738 segments · language en · Watch on YouTube

Full transcript

  1. 0:01Everyone is talking about AI agents
  2. 0:02these days.
  3. 0:15But the reality is that most of us are
  4. 0:17still using AI the same way we did a
  5. 0:19year ago. For example, we type a
  6. 0:23question into ChatGPT or Claude, get an
  7. 0:25answer, copy it, and use it elsewhere,
  8. 0:28but we can use AI for so much more now.
  9. 0:32By using agents, we can do much, much
  10. 0:34more interesting things. See, a chatbot
  11. 0:38only lives inside that chat window. You
  12. 0:42ask it a question, and it gives you an
  13. 0:44answer. It has no access to other tools
  14. 0:46. In contrast, an agent has that same
  15. 0:50brain, but imagine it also has hands
  16. 0:52and feet. It has access to a set of
  17. 0:57tools, access to your files, access to
  18. 0:59the web so it can search, it has some
  19. 1:01memory, specific goals, and infinite
  20. 1:03tools you can connect to it. In this
  21. 1:12video, I want to show step-by-step what
  22. 1:14an agent is capable of, what its
  23. 1:15internal anatomy is, what role each
  24. 1:17part of an agent plays in its
  25. 1:19performance, and knowing this will help
  26. 1:21you understand where you can use an
  27. 1:22agent in your work, what you need to
  28. 1:24build one, and this is the most
  29. 1:26important step for optimally
  30. 1:27integrating AI into your work and
  31. 1:29projects. Well, look, the foundation of
  32. 1:37all these tools is one thing: a Large
  33. 1:39Language Model, or an LLM, which acts
  34. 1:41as the brain of this system; like
  35. 1:44ChatGPT, Gemini, Claude, and all of
  36. 1:46those. They are applications built on
  37. 1:50top of a language model, and they work
  38. 1:51simply. You give it an input, and the
  39. 1:55model gives you an output based on the
  40. 1:57data it was trained on. For example,
  41. 2:01you ask it to write an email on a topic
  42. 2:03you choose; your prompt is the input,
  43. 2:05and the email text you receive, which
  44. 2:07is usually more polite than what we
  45. 2:08write ourselves, is the output of this
  46. 2:10language model. Now, ask that same
  47. 2:14model when my next meeting is, and it
  48. 2:16cannot answer. Why? Because it doesn't
  49. 2:18have access to your data. It doesn't
  50. 2:21have access to your calendar, and this
  51. 2:23actually highlights two important
  52. 2:24characteristics of these language
  53. 2:26models. Two features that are very
  54. 2:29important for understanding agents. One
  55. 2:32is that these models are indeed trained
  56. 2:34on a massive amount of data; from the
  57. 2:36whole world, from all sources. But they
  58. 2:39know nothing about your private
  59. 2:41information, your calendar, your emails
  60. 2:43, or your internal company documents.
  61. 2:46It’s like a brain that knows all of
  62. 2:48human knowledge, but it doesn't know
  63. 2:50you specifically, nor does it know the
  64. 2:51details of your life. Secondly, these
  65. 2:56language models are passive; meaning
  66. 2:58they sit and wait for you to give them
  67. 3:00a prompt, ask a question, so they can
  68. 3:02respond and provide the text answer to
  69. 3:04you. It doesn't start anything on its
  70. 3:07own, and because it lacks access to
  71. 3:09tools, it can't perform any actual
  72. 3:11tasks. As a very simple example,
  73. 3:14imagine a highly professional chef
  74. 3:16sitting in an empty room. Now imagine
  75. 3:21that same chef running the kitchen in
  76. 3:23the best restaurant in town. Why?
  77. 3:26Because they have access to various
  78. 3:28tools like the stove, fresh ingredients
  79. 3:30in the fridge, and professional cooking
  80. 3:32utensils, allowing them to prepare the
  81. 3:34best dishes. But you can only go to the
  82. 3:36chef sitting in the empty room and ask
  83. 3:38questions like, "What should I do for
  84. 3:40this dish?""What do you think about
  85. 3:42this?" Consequently, this is the
  86. 3:45difference between an AI agent and a
  87. 3:47regular chatbot like ChatGPT or Claude.
  88. 3:50Of course, Claude and ChatGPT have
  89. 3:52advanced significantly now and their
  90. 3:54agentic capabilities have increased,
  91. 3:56but we are using their chat mode as an
  92. 3:58example. Before we get to agents
  93. 4:00themselves, let's have a general
  94. 4:02classification together. The use of AI
  95. 4:07can be categorized into 4 levels, and
  96. 4:09most people, without knowing higher
  97. 4:11levels exist, remain at the first level
  98. 4:13, which is using chatbots. So, level
  99. 4:17one is chat. Meaning you ask a question
  100. 4:20and get an answer. This is very useful
  101. 4:23and has already transformed our lives
  102. 4:25significantly. But you have to do the
  103. 4:27rest of the work yourself. You have to
  104. 4:29make the decisions. When to ask a
  105. 4:31question? What to ask, and then you
  106. 4:33have to copy the text, take it
  107. 4:34somewhere else, and do whatever you
  108. 4:36need to do; you have to handle the rest
  109. 4:38yourself. Level two is AI tools.
  110. 4:41Meaning tools that can perform specific
  111. 4:44tasks, for example. For instance, you
  112. 4:47see an AI tool dedicated to generating
  113. 4:49images. One tool is dedicated to
  114. 4:52creating slides. One tool is dedicated
  115. 4:55to analyzing tax documents, and
  116. 4:57ultimately, they deliver a final
  117. 4:59product. But you are still the one
  118. 5:02behind the wheel. Meaning you decide
  119. 5:04yourself, "I need to go use this tool
  120. 5:06right now.""Delegate this task to it.""
  121. 5:10Get this output from it." What is level
  122. 5:12three? AI workflows; this is where
  123. 5:14tasks are chained together. Meaning you
  124. 5:20define a series of specific tasks you
  125. 5:22want performed in a sequence—that is,
  126. 5:24"do this, then do that, then do this"
  127. 5:26—within a pipeline, and from then on,
  128. 5:28it automatically executes them one
  129. 5:30after another. For example, assume that
  130. 5:36every morning you take links to several
  131. 5:38news items from LinkedIn, ask a model
  132. 5:40to summarize them, and then ask other
  133. 5:42models to write a LinkedIn post based
  134. 5:44on that summary. This is an example of
  135. 5:50a workflow where you perform a series
  136. 5:52of different tasks in a chain, each of
  137. 5:53which might even require AI. Meaning,
  138. 5:59you can build such a workflow using
  139. 6:01tools like n8n or Make and set it to
  140. 6:03run for you every day at 8 AM. The
  141. 6:08important feature of a workflow is that
  142. 6:09it can only follow a path that a human
  143. 6:11has predefined for it. For instance, if
  144. 6:16you ask a calendar workflow about the
  145. 6:18weather, well, it cannot answer.
  146. 6:21Because the path you defined only goes
  147. 6:22through the calendar. It has no access
  148. 6:25to anything else. In short, when a
  149. 6:27human has decided what the workflow
  150. 6:29path should be, it is no longer an
  151. 6:31agent. It is a workflow. Now, before we
  152. 6:35reach the fourth level, or agents,
  153. 6:37let's talk about the concept of RAG.
  154. 6:40RAG is an abbreviation for
  155. 6:42Retrieval-Augmented Generation. What is
  156. 6:44its general concept? Since we are going
  157. 6:46through this video in a summarized
  158. 6:48manner. The general concept is that you
  159. 6:52provide a database or an information
  160. 6:54source to the AI you are using. This
  161. 7:02means you can provide specific
  162. 7:04information—in the form of PDFs,
  163. 7:06databases, tables, images, or anything
  164. 7:08you can think of—in its own specific
  165. 7:10format to the language model and ask it
  166. 7:12to refer to that database before
  167. 7:14answering your question, and provide an
  168. 7:16answer based on that information. Now
  169. 7:23we reach level 4, which is agents. The
  170. 7:28only major change that needs to happen
  171. 7:30for a workflow to become an agent is
  172. 7:32that the human decision-maker must be
  173. 7:33replaced by a language model. With an
  174. 7:38agent, you no longer give it a
  175. 7:40step-by-step workflow path. You entrust
  176. 7:43it with a goal. Meaning, you don't say
  177. 7:45do this first and then do that. You say
  178. 7:47, "I want this result.""These are the
  179. 7:49tools at your disposal." The agent
  180. 7:52itself then finds the path to reach
  181. 7:53that goal. Structurally speaking, an
  182. 7:57agent is just a language model
  183. 7:58connected to four things. First, the
  184. 8:01language model itself, which is its
  185. 8:03reasoning engine or its brain. Second,
  186. 8:07the tools we provide, such as the
  187. 8:09terminal, browser, file systems it can
  188. 8:11access, and APIs it has access to—
  189. 8:13these are like the system's hands with
  190. 8:16which it can perform tasks. Third,
  191. 8:19memory. This is very important; it's
  192. 8:23like information the agent reads at the
  193. 8:25start of each session so it doesn't
  194. 8:27start from scratch—a kind of memory
  195. 8:29of previous tasks performed and
  196. 8:31previous conversations had, so it can
  197. 8:33be assigned longer tasks. And fourth:
  198. 8:38the goal; which means specifying what
  199. 8:40result you want to achieve, with a
  200. 8:42clear definition of that result or
  201. 8:43destination you provide it. One
  202. 8:48important point that brings agents to
  203. 8:50life is, in a way, the loop. See, every
  204. 8:54agent on every platform works with a
  205. 8:56three-step loop. Meaning the agent
  206. 9:01observes the current state: what files
  207. 9:03exist, what image a webpage is showing,
  208. 9:05or what result the last command
  209. 9:07returned. Then "thought"; meaning based
  210. 9:13on what you just said and the current
  211. 9:15conditions, it thinks and decides what
  212. 9:17the right next step is. And "action";
  213. 9:24changing something using one of the
  214. 9:26tools at its disposal; should it write
  215. 9:28a file, perform a web search, or
  216. 9:30execute a code command? And this cycle
  217. 9:36repeats until it reaches the result and
  218. 9:37the goal we initially set. If you see
  219. 9:43this term "ReAct" written with a
  220. 9:45capital A and a lowercase a, that is
  221. 9:47the "React" programming language; but
  222. 9:49if you see it with a capital A, it is a
  223. 9:51combination of the two words "Reason"
  224. 9:53and "Act," because this is the
  225. 9:54combination that agents perform. They
  226. 9:58reason, they take action. They reason,
  227. 10:01they take action, and this is the most
  228. 10:03common framework for agents. Well, so
  229. 10:06far we have seen the skeleton of an
  230. 10:08agent and what it is composed of. Let's
  231. 10:12dive a little deeper into some of its
  232. 10:14parts. The memory section is an
  233. 10:17important part of agents. We said that
  234. 10:20a language model doesn't know you
  235. 10:21specifically. But agents usually have
  236. 10:24four types of memory. When I speak of a
  237. 10:27language model, I am not specifically
  238. 10:29talking about ChatGPT or Claude;
  239. 10:31because ChatGPT and Claude are not just
  240. 10:33raw language models. They also have a
  241. 10:36memory of you; if you notice, they
  242. 10:38gradually learn and understand more
  243. 10:40things about you. When I talk about a
  244. 10:43language model, the raw language model
  245. 10:46is ChatGPT. If you have used ChatGPT,
  246. 10:50Claude, and similar tools, these also
  247. 10:52have some agentic capabilities added to
  248. 10:54them; especially Claude, if you have
  249. 10:56used its software, it has a very, very
  250. 10:59strong agentic framework. Now let's go
  251. 11:03back and talk about the types of memory
  252. 11:05so it becomes clearer for you. The
  253. 11:10first memory is working memory; every
  254. 11:12time you talk to the agent, your
  255. 11:14question, the history of this
  256. 11:15particular conversation, all the Q&A
  257. 11:17between you and the chatbot, and the
  258. 11:19instructions given to the system
  259. 11:21beforehand are placed in a temporary
  260. 11:23package; similar to computer RAM. This
  261. 11:28is a volatile memory. Meaning when the
  262. 11:30conversation ends, these chats are
  263. 11:32cleared. Ordinary chatbots only have
  264. 11:34this. For example, if you go to a site
  265. 11:37and talk to its chatbot support, close
  266. 11:39the page, and come back, everything
  267. 11:41starts from scratch because that memory
  268. 11:43only held the memory of that particular
  269. 11:45chat. The second type is procedural
  270. 11:48memory. This memory actually tells the
  271. 11:51agent how to behave. For example, a set
  272. 11:54of instructions, rules, or workflows is
  273. 11:57given to it, usually as a text file
  274. 11:59with an extension. It is saved as MD,
  275. 12:03which is called a skill. You might hear
  276. 12:06this term a lot these days. This skill
  277. 12:09is actually a text file in the format
  278. 12:11of that same Markdown or. MD file,
  279. 12:14teaching the agent how to perform a
  280. 12:16specific task. The third type of memory
  281. 12:19is semantic memory. A series of stable
  282. 12:22facts that we want the agent to always
  283. 12:24know. For example, who are you? What is
  284. 12:28your product? What questions do
  285. 12:29customers usually ask? Exactly the
  286. 12:33things that a language model doesn't
  287. 12:35know about you, but you want your agent
  288. 12:36to know. And the fourth type of memory
  289. 12:43is episodic; meaning a time-stamped
  290. 12:45history of information; meaning that
  291. 12:47what was said at any given time and
  292. 12:49what happened are recorded, and the
  293. 12:50agent is aware of them, like a logbook
  294. 12:52where events are recorded by time. For
  295. 12:59instance, suppose you have a problem
  296. 13:00and you tell your agent. Your agent
  297. 13:03solves it and records it somewhere. It
  298. 13:08says, for example, on such and such a
  299. 13:10day, this issue was resolved, and later
  300. 13:12if you refer to it, since your agent
  301. 13:14has this memory, it says this problem
  302. 13:15occurred on such a date and was solved,
  303. 13:17and this can be implemented for
  304. 13:19everything. For all types of dated
  305. 13:24event logs; for example, if it is
  306. 13:26inventory management, or an educational
  307. 13:28topic, or a legal matter, and well, you
  308. 13:30can imagine how many places it is used.
  309. 13:37Before moving to the next section, let
  310. 13:40me mention one point again: if you need
  311. 13:42consultation regarding integrating AI
  312. 13:44into your workflow, your projects,
  313. 13:45building agents, or new AI projects,
  314. 13:47the way to contact me for consultation
  315. 13:49is in the description of this video.
  316. 13:56You can act through my website and
  317. 13:58schedule a consultation session. Now,
  318. 14:01these memories are stored in a series
  319. 14:03of databases. Whenever the agent needs
  320. 14:08to find one of these pieces of
  321. 14:10information, it refers to the part that
  322. 14:11it knows is programmed for that type of
  323. 14:13information. For example, among the
  324. 14:191,000 conversations you've had with the
  325. 14:21agent, it can find 20 specific
  326. 14:23conversations about a certain topic,
  327. 14:25and it matches the meaning of that text
  328. 14:27with what it's looking for; it doesn't
  329. 14:29necessarily have to search and find the
  330. 14:31exact same words. Now, an important
  331. 14:37point is that from a certain point on,
  332. 14:39the text of the conversations between
  333. 14:41the person and the agent grows; meaning
  334. 14:43it might have 1,000 lines of
  335. 14:45conversation, and as this volume gets
  336. 14:47larger and larger, it becomes expensive
  337. 14:49and the agent's accuracy decreases. As
  338. 14:57a result, what they do after a certain
  339. 14:59point—for example, once a month
  340. 15:01passes—is summarize all the previous
  341. 15:03month's conversations (it's called "
  342. 15:05distilling," I think you could also
  343. 15:07call it "refining"), extract the
  344. 15:09essence, derive a few summary sentences
  345. 15:11, pull out the key points, and store
  346. 15:13those specific sentences in their
  347. 15:15semantic memory. This is what they call
  348. 15:20memory consolidation. They consolidate
  349. 15:23certain things into memory, which in
  350. 15:25turn makes the language model run more
  351. 15:27cost-effectively. Now, these agents
  352. 15:30have a kind of memory file. You might
  353. 15:34have seen this if you have worked with
  354. 15:35agents. For some tools, it is called
  355. 15:38cloud.md. For some tools, it is called.
  356. 15:42md. There are actually a set of rules
  357. 15:47in this file that the agent reads every
  358. 15:49time it wants to start talking or
  359. 15:50working, so it is aware of them before
  360. 15:52it begins. For a very simple example,
  361. 15:57suppose you tell an agent to, say,
  362. 15:59build a website for you. Every time it
  363. 16:02builds a website, it puts an emoji next
  364. 16:04to every sentence, and you have to keep
  365. 16:06deleting them. But if you ask it to
  366. 16:11record this rule in its memory file—
  367. 16:13that in the text you put on the site or
  368. 16:15in the comments, never use emojis—it
  369. 16:17will never repeat that mistake again,
  370. 16:19or even more professionally, if you
  371. 16:21manually know where that cloud.md file
  372. 16:23is. Where is the md file? You can add a
  373. 16:29sentence to it telling it not to do
  374. 16:31that anymore or to do it a certain way,
  375. 16:33and the agent then learns to structure
  376. 16:34its work based on the rules you have
  377. 16:36defined. Now, let's talk a bit about
  378. 16:41some terms that have recently become
  379. 16:43very popular regarding agents or that
  380. 16:45we hear quite often. One of these terms
  381. 16:49is "harness." Literally, harness means
  382. 16:53a bridle like a horse's harness, and
  383. 16:55here it is used as a model harness. If
  384. 17:01you imagine a language model as a
  385. 17:03powerful horse, it might run anywhere
  386. 17:04without control, but with a harness,
  387. 17:06you can control it to do exactly what
  388. 17:08you want. Now, why is control necessary
  389. 17:12? Look, at their core, language models
  390. 17:16are fundamentally predicting the
  391. 17:17probability of the next word. This is
  392. 17:21how they write a long text. Wherever
  393. 17:23there is probability, there is also the
  394. 17:24possibility of error. In fact, where
  395. 17:28you are just chatting, it doesn't
  396. 17:29matter much, but where the task you are
  397. 17:31performing is important—you might
  398. 17:33want your agent to modify files, make
  399. 17:35certain decisions, or perform specific
  400. 17:37actions—the probability of this error
  401. 17:39must be minimized as much as possible.
  402. 17:44Harnessing is essentially all the
  403. 17:46guardrails we place around a model to
  404. 17:47ensure it behaves exactly the way we
  405. 17:49want. Harnessing is a broad topic;
  406. 17:53I’ll probably make a separate video
  407. 17:55for it. But for now, just know that
  408. 18:01harnessing can include specifying what
  409. 18:02the model should or shouldn't do in its
  410. 18:04system prompt, exercising caution when
  411. 18:06defining tools, the instructions you
  412. 18:08provide, and managing the loops that
  413. 18:10repeat until the desired result is
  414. 18:12achieved. Keep this general idea in
  415. 18:18mind; I’ll try to discuss it more
  416. 18:20later. If you're interested in loops,
  417. 18:25that’s also an important term
  418. 18:26currently discussed regarding agents.
  419. 18:33Engineering these loops and setting
  420. 18:35stop guardrails are also key topics in
  421. 18:36building and using agents. I’ve
  422. 18:42mentioned before that an agent rotates
  423. 18:43through an observation, thought, and
  424. 18:45action loop; but the important question
  425. 18:47is, when should this loop stop? The
  426. 18:50simple answer is when it reaches its
  427. 18:52goal. But how does the agent know it
  428. 18:54has reached its goal? It might call a
  429. 18:58tool 20 times, fail to reach the goal,
  430. 19:00and then be stuck calling that tool
  431. 19:01infinitely. Consequently, it needs to
  432. 19:07know at some point how "good" is enough
  433. 19:09—that is, what result is sufficient,
  434. 19:11or if it can't reach a conclusion, we
  435. 19:13need to define for it when to stop the
  436. 19:15loop. For example, you tell your agent:
  437. 19:21"See what customers are complaining
  438. 19:23about; if anyone has a refund request
  439. 19:25that hasn't been processed, follow up
  440. 19:27on it." The agent first pulls data from
  441. 19:32the customer management system; for
  442. 19:34instance, suppose there were 30
  443. 19:35complaints in the last 2 months, 12
  444. 19:37refunds were processed, and 8 remain.
  445. 19:41Then it thinks to itself that it was
  446. 19:43asked to follow up on the remaining
  447. 19:45cases; should I follow up with these 8
  448. 19:46people now? Should I schedule a meeting
  449. 19:49with them? Should I send them an email?
  450. 19:52Should I start the refund process
  451. 19:54myself? And well, this is an open-ended
  452. 19:56scenario. The important point is that
  453. 19:59for every process, you must anticipate
  454. 20:01and define an endpoint. If you want to
  455. 20:07prevent errors, a good agent is one
  456. 20:08that, when you ask it for a task, asks
  457. 20:10you during the planning phase: "What is
  458. 20:12the end goal in your opinion?" Should I
  459. 20:17process the refund myself, or just
  460. 20:19deliver the list? Here, you are the one
  461. 20:22defining the loop's termination
  462. 20:24condition. This is the stage agents are
  463. 20:26currently in. Certainly, within a year,
  464. 20:30they will be able to make these
  465. 20:31decisions much more easily, and in a
  466. 20:33way, they will give you choices based
  467. 20:35on your preferences. I'll give you
  468. 20:38another important example of loop
  469. 20:39engineering. Look, consider a situation
  470. 20:44where an agent might ask you a question
  471. 20:46to plan or request permission to access
  472. 20:48something, and you've left the agent to
  473. 20:50do the job while you've gone off; this
  474. 20:52might have happened to you; then, for
  475. 20:54instance, you return after half an hour
  476. 20:56, happy to see the result, only to find
  477. 20:58the agent has just asked, "Can I have
  478. 20:59access to this file or not?" And it's
  479. 21:04still waiting for file access
  480. 21:05permission; essentially, half an hour
  481. 21:07or several hours of your time have been
  482. 21:09wasted. How, how can you prevent this?
  483. 21:14In the same harness or file, define a
  484. 21:16rule for it: whenever you are waiting
  485. 21:18for my permission for something, send
  486. 21:20me a notification on Telegram. You know
  487. 21:26, these are things that can be defined
  488. 21:27to optimize the agent's running process
  489. 21:29as much as possible. Well, a question
  490. 21:33that arises is, we build an agent. It
  491. 21:36has memory, tools, and its loop is also
  492. 21:38refined. How do we know this agent
  493. 21:41works well? Here we reach the topic of
  494. 21:44LLMOps, similar to DevOps. The "Ops"
  495. 21:48part stands for operations, meaning
  496. 21:50language model operations. In fact,
  497. 21:56evaluation of language models starts
  498. 21:58with tracking; meaning everything from
  499. 21:59the moment the user asks a question to
  500. 22:01the answers the agent gives, the tools
  501. 22:03it calls, and the thoughts it has, must
  502. 22:05be recorded. How many times each tool
  503. 22:11was called, how long each step took,
  504. 22:13how many tokens were consumed—these
  505. 22:15must be recorded. The second step is
  506. 22:18evaluation. Evaluating to see if this
  507. 22:21execution was good, if it was done
  508. 22:22correctly. Interestingly, you can even
  509. 22:27use another language model to evaluate
  510. 22:29the quality of an agent's work. Meaning
  511. 22:34you use another language model to
  512. 22:36evaluate whether this agent performed
  513. 22:38its task well relative to what was
  514. 22:39requested or not. Or you yourself can
  515. 22:43objectively decide whether it did a
  516. 22:45good job or not. For example, if you
  517. 22:48had asked it to follow up on customer
  518. 22:50work, did it do it the way you wanted
  519. 22:51or not? If it was supposed to send an
  520. 22:53email, did it send it or not? The third
  521. 22:55step is diagnosis and correction.
  522. 22:57Meaning, for example, you detect that
  523. 22:59this agent took too long to finish its
  524. 23:01task. Consequently, you must diagnose
  525. 23:04the reason for this excessive duration
  526. 23:06there. Has the working memory become
  527. 23:08too large? Maybe, for a simple question
  528. 23:11, it's searching the entire system
  529. 23:13memory. Perhaps it only needs access to
  530. 23:16a small portion of the memory. That's
  531. 23:18where the memory system needs to be
  532. 23:20engineered and fixed. Or maybe a
  533. 23:22question doesn't require memory
  534. 23:23retrieval at all. For example, what is
  535. 23:26the capital of France? The model no
  536. 23:28longer needs to check its memory; the
  537. 23:30LLM itself already holds that
  538. 23:31information. Or perhaps the prompt or
  539. 23:35instructions you give your agent, for
  540. 23:37example, saying: "Don't overcomplicate
  541. 23:39it, just use a simple approach," or in
  542. 23:41cases where you need more complex
  543. 23:42analysis, you should tell it to "think
  544. 23:44more deeply." These are the parts where
  545. 23:48, after reviewing, you identify a
  546. 23:50series of flaws and correct them. So,
  547. 23:53the complete LLM-Ops cycle is tracking,
  548. 23:56evaluation, identification, and
  549. 23:58correction. If your agent is built well
  550. 24:02, it's a system that can even improve
  551. 24:04itself over time; this is where strong,
  552. 24:06real agents are separated from weak
  553. 24:08ones. Agents that can improve
  554. 24:13themselves over time. Now, this is
  555. 24:17actually one of the most practical
  556. 24:18parts of the video that doesn't get
  557. 24:20much attention. The way you write
  558. 24:25prompts for an agent should differ from
  559. 24:26the way you write prompts for a chatbot
  560. 24:28. When you write a prompt for a chatbot
  561. 24:33, you are essentially describing what
  562. 24:34you want from the chatbot. But when you
  563. 24:39want to give a prompt to your agent,
  564. 24:41it's as if you are signing a contract
  565. 24:42with it. This means you must provide
  566. 24:48precise instructions based on which it
  567. 24:50is obligated to perform the work and
  568. 24:51deliver it to you. This is very
  569. 24:55important. The more precise the
  570. 24:59instructions and the more clearly you
  571. 25:01define potential cases—telling it "in
  572. 25:03this scenario do this, and in that
  573. 25:04scenario do that"—the more smoothly
  574. 25:06your agent will operate. For instance,
  575. 25:12let me give an example: you tell a
  576. 25:14chatbot to build a website page for my
  577. 25:16new product. Well, you’ve said
  578. 25:20something general, and it's also very
  579. 25:21vague. It can fill in the blanks you
  580. 25:25haven't fully specified however it
  581. 25:26wants and get quite creative; it
  582. 25:28essentially just gives you the code.
  583. 25:33But give the same text to an agent, and
  584. 25:36well, the agent has tools, loops,
  585. 25:38access to various resources, and
  586. 25:39authority; it starts building a
  587. 25:41framework it deems appropriate, which
  588. 25:43might take half an hour. Now, depending
  589. 25:49on the task you ask of it, one job
  590. 25:51might take 5 minutes, another might
  591. 25:52take an hour or two of your time, and
  592. 25:54in the end, it might turn out to be
  593. 25:56something you never wanted at all; it
  594. 25:58depends on how specifically you know
  595. 26:00what you want and it will cost you. As
  596. 26:04a result, it is very important that
  597. 26:06your prompt is like a contract; specify
  598. 26:08exactly what things it can use. What
  599. 26:12systems should it use? What frameworks
  600. 26:15should it not use? Ultimately, the more
  601. 26:19precisely you write this contract, the
  602. 26:21more useful an agent with fewer errors
  603. 26:23you can have. The issue is clearly not
  604. 26:27just about longer prompts. The prompt
  605. 26:30should be structured. A structured
  606. 26:32prompt has four parts. First, its goal
  607. 26:36must be completely clear; not just what
  608. 26:38task it should do, but what "finished"
  609. 26:40looks like exactly from your
  610. 26:42perspective. What do you want the final
  611. 26:45output to look like? The second point:
  612. 26:47constraints, what it is not allowed to
  613. 26:50do; for example, "do not install new
  614. 26:51packages without my permission," or "do
  615. 26:54not make changes to the database
  616. 26:55without my approval." Whenever
  617. 27:01something wasn't on your list of
  618. 27:03constraints and you later see your
  619. 27:04agent performed that action, you can
  620. 27:06add those things—which you think of
  621. 27:08later—to the agent's constraint list;
  622. 27:10in a way, much of this comes with
  623. 27:12experience, and it can vary across
  624. 27:13different tasks. And another point:
  625. 27:18what is the exact format or structure
  626. 27:20of the output? For instance, the output
  627. 27:22you want delivered to you, do you want
  628. 27:24it in PDF format? Do you want it to be
  629. 27:26an HTML file? Do you want it saved in a
  630. 27:28database? You must specify these things
  631. 27:32completely and think about the points
  632. 27:34where it might fail; meaning, where
  633. 27:35might it get stuck? What should it do
  634. 27:38when it gets stuck? Should it stop
  635. 27:40completely when it gets stuck? Should
  636. 27:43it ask you for input? Should it send
  637. 27:45you a notification on Telegram, like in
  638. 27:47the previous example? These are things
  639. 27:50you can define, and it is very
  640. 27:51necessary to define them for agents. So
  641. 27:54, knowing this, where should we start
  642. 27:56building an agent? Look, depending on
  643. 27:59who you are and what you want from an
  644. 28:00agent, there are different paths. I
  645. 28:04don't intend to teach it in this video,
  646. 28:05but I will mention the different modes.
  647. 28:09Write in the comments which one is more
  648. 28:10interesting to you so I can teach it in
  649. 28:12the next video. Whichever gets more
  650. 28:14votes, I will definitely make a
  651. 28:16tutorial video for it. If you are
  652. 28:20someone who wants an all-purpose agent
  653. 28:22on your own desktop computer, tools
  654. 28:24like Anthropic's Claude Code or
  655. 28:25OpenAI's Codex are the best options for
  656. 28:27you. Meaning they are agents themselves
  657. 28:32. They have access to many tools
  658. 28:34themselves. You can also add many tools
  659. 28:37to them later yourself, and they can
  660. 28:38perform many tasks for you. From
  661. 28:41organizing files and extracting PDFs to
  662. 28:43building things and creating apps, they
  663. 28:46can do many of these tasks. But if you
  664. 28:51are into building workflows, for
  665. 28:52example, you can use tools like n8n or
  666. 28:54Make. I have taught n8n in two of my
  667. 29:00videos. In n8n, you can build more than
  668. 29:03just workflows. You can build agents.
  669. 29:05If you haven't seen those videos of
  670. 29:07mine, definitely go watch them. I will
  671. 29:08put the link in the description. If you
  672. 29:12want to build a personal assistant
  673. 29:14without any coding that, for example,
  674. 29:16checks your emails or replies on
  675. 29:17WhatsApp. Platforms like Flowise and
  676. 29:23Automa. Can be useful for you. The
  677. 29:28Automa platform. I also talked about it
  678. 29:31in my previous video. If you're
  679. 29:33interested in Basecamp 4, let me know
  680. 29:34and I'll make a video about it. But if
  681. 29:38you're a developer and want to dive
  682. 29:40deep into the technical side to build
  683. 29:42exactly what you need, and of course
  684. 29:44build much more powerful things, you
  685. 29:46can create your own agents using
  686. 29:47frameworks like LangChain or Pydantic
  687. 29:49AI. Define your own local models. Use
  688. 29:54local models. Assemble the components
  689. 29:57yourself and, in fact, keep control
  690. 29:59over all the details. Especially
  691. 30:03Pydantic AI, which in my opinion is
  692. 30:05very, very strong for agentic systems
  693. 30:07right now. Where LangChain and
  694. 30:10LangGraph show certain limitations,
  695. 30:12Pydantic AI is quickly resolving them
  696. 30:14and moving forward. Now, neither of
  697. 30:17these methods is better than the other.
  698. 30:20It just depends on what you need from
  699. 30:22building an agent, so choose based on
  700. 30:24that. Which one of these methods to use
  701. 30:27. Let's do a short recap. Three
  702. 30:31important things we learned in this
  703. 30:32video. First, about agent architecture:
  704. 30:38an agent is essentially a language
  705. 30:40model that, besides tools, also has
  706. 30:42memory and a goal, which are connected
  707. 30:44by a loop of observation, thought, and
  708. 30:46action. Now, the next time someone says
  709. 30:52"AI agent," you'll know exactly what
  710. 30:54they're talking about. Second, the
  711. 30:57discussion on memory was important;
  712. 30:59memory can be a simple text file that
  713. 31:01you place alongside your agent. It
  714. 31:06could be data that is continuously
  715. 31:07recorded in your database, which your
  716. 31:09agent has access to. And third, the
  717. 31:17prompt contract, which is very
  718. 31:18important because it defines the
  719. 31:20agent's goal, its constraints, the
  720. 31:22output format, and what it should do in
  721. 31:24case of failure or errors. These were
  722. 31:30the three important things about agents
  723. 31:32. I hope you found it interesting. I
  724. 31:36would like to continue this series on
  725. 31:37agents. To teach more about it. And the
  726. 31:41various tools available for building
  727. 31:43agents. Be sure to write in the
  728. 31:46comments exactly what interests you for
  729. 31:47future videos. If you need my
  730. 31:52consultation regarding your projects,
  731. 31:54building agents, or creating automation
  732. 31:56workflows for your company or work, the
  733. 31:57link to contact me is in the
  734. 31:59description. You can definitely use
  735. 32:04that link through my site to get in
  736. 32:05touch with me, and as always, thank you
  737. 32:07for watching. For now, until the next
  738. 32:11video, goodbye.

About this transcript

This page contains the full transcript of همه‌چیز دربارهٔ AI Agent — از مدل زبانی تا ایجنت خودمختار by Maryam Sadeghi, generated from the public captions YouTube serves with the video. The transcript has 4,904 words across 738 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.