8 сентября 2026 г. — Transcript
Full transcript
- 0:00Hey everyone, this is Sean. Today let's
- 0:01think about how to design an AI agent
- 0:03system like a pro. It doesn't matter if
- 0:05you're technical or not. If you're a
- 0:06product manager, designer, data
- 0:07scientist or computer scientist, does
- 0:09not matter. The goal here really is to
- 0:11understand how do we think holistically
- 0:13about you know starting from the front
- 0:15end all the way to the back end database
- 0:17and setting up AI agents so that your
- 0:19app will be smarter than just the
- 0:21traditional SAS product. So today's
- 0:23example is about how do we design this
- 0:25agent system for the e-commerce direct
- 0:27to customer brands D2C brands to manage
- 0:30their customer support. So as we know
- 0:32that customer support is a field that
- 0:33traditionally has a lot of human
- 0:35involvement when it comes to things like
- 0:37return products, exchange products or
- 0:39you know asking about the status of
- 0:41where my product is when it's being
- 0:42shipped. And u for a lot of e-commerce
- 0:45company they probably cannot afford to
- 0:47hire a bunch of like call centers or
- 0:48message replying agents. So this would
- 0:51be a perfect use case for them to you
- 0:53know increase their response rate as
- 0:55well as you know um improve the
- 0:57satisfaction for their customers. So now
- 0:59let's jump in and think about how do we
- 1:01think through this step by step. Okay.
- 1:03So the first thing is that our goal
- 1:04today is that we want to make sure that
- 1:08we will have over 70%
- 1:11automation in this customer support
- 1:13flow. Which means that the other 30%
- 1:15probably be like we're going to loop the
- 1:17humans back into the conversation when a
- 1:21e-commerce site is talking to a
- 1:22customer, right? And also we want to
- 1:24make sure that the customer satisfaction
- 1:26rate sees greater than 4.5 out of five,
- 1:30right? And last but not least, we want
- 1:31to make sure that the response with the
- 1:3450% of the percentile of the customers
- 1:36is below 1 second. And then 95% of your
- 1:40customers will get a response under 2.5
- 1:42seconds. Why is this important? This is
- 1:44important because if the response time
- 1:47is too long, it really affects your
- 1:49satisfaction from your customers. Okay,
- 1:52so this is our goal. And um let's also
- 1:54think about what is our scope today. Our
- 1:57scope is very simple. We want to focus
- 2:00on returns, exchanges,
- 2:04and where is my order? Okay, so let me
- 2:08just delete this flow and then let's
- 2:10walk through it together. So firstly,
- 2:12let's think about the rules of this
- 2:14scope.
- 2:16In order to return to the customer, we
- 2:18need to make sure that the products for
- 2:21example will only be returned if it's
- 2:23bought over under 30 days before. It
- 2:25must be returned in good conditions,
- 2:26right? Otherwise, the agents probably
- 2:28will just allow any products to be
- 2:30returned at any time, which is not what
- 2:31we want, right? The second rule we think
- 2:33about is um how do we think about the
- 2:35refund policy, right? In this case, most
- 2:38companies probably have a refund policy.
- 2:40for example, if it's under $50 and if
- 2:42they're using a human call center uh as
- 2:44a service, then they will automatically
- 2:47let it to um return the product. But if
- 2:50it's over or equal to $50, then they
- 2:52probably need to loop in their manager.
- 2:53We want to mimic the exact user behavior
- 2:55here. The third one is called exchange
- 2:58policy. So if a customer want to
- 3:00exchange the product, they also need to
- 3:02make sure a few things. um it needs to
- 3:05be in good conditions and certain
- 3:06categories of products cannot be
- 3:08returned because for example if it's
- 3:10food and you already open it you cannot
- 3:11really return it right
- 3:14and lastly that um how do we route back
- 3:16to a human right and there could be all
- 3:18sorts of scenarios here um it could be
- 3:21because the user directly asked that
- 3:23they want to talk to a human or we have
- 3:25detected that there's some very negative
- 3:28emotions in the chat so that we must uh
- 3:31involve a human being to basically
- 3:34provide emotional support for the
- 3:36customer. Okay. Um from a backend and
- 3:39database perspective, there are a few
- 3:41things that we should consider in this
- 3:43case. Let me just paste it right in. The
- 3:45first thing is that when and how much is
- 3:48the peak for requests per second and
- 3:52about the seasonality as well. Imagine
- 3:53this is Black Friday or Christmas time.
- 3:55Then probably there will be a lot more
- 3:56people doing online shopping versus the
- 3:59other period of time of the year. And
- 4:01also we need to think about how are we
- 4:03going to store all the chat history
- 4:06um information about the transactions
- 4:09about return and logistics. Do we store
- 4:11it in a relational database or do we
- 4:14store it in a vector database? So here's
- 4:16a concept we need to introduce for AI
- 4:18agents which is vector database. What it
- 4:20really means is that traditionally we
- 4:24basically define data tables in imagine
- 4:27you have an Excel sheet there are rows
- 4:28and columns. every column means a
- 4:30different feature. Every row means like
- 4:32a different entry, right? Um but in
- 4:36order to search some of the say
- 4:37unstructured data like the return policy
- 4:40or refund policy, these kind of things
- 4:42are not going to be stored in
- 4:44traditional like table format. They're
- 4:45probably a paragraph or PDF, right? So
- 4:48the way to do is that we can store this
- 4:50information in a vector database in
- 4:53which we will embed each word or each
- 4:56you know paragraphs into uh vectors.
- 5:00Right? So what we mean by that is that
- 5:01we're going to turn these words into
- 5:04numbers that are in high dimensions. If
- 5:06you're really not from a technical
- 5:08background, just think of it as like we
- 5:09got to digitize something. Like if
- 5:11you're looking at a photo on your
- 5:13iPhone, it will be stored in digits and
- 5:16then when they're stored in digits,
- 5:18you'll be able to search them by
- 5:19calculating some similarity between
- 5:21photos. Similarly, here we're allowing
- 5:23it to calculate similarities between uh
- 5:26vectors. Right? So given this current
- 5:29scope, the goal, the rule, the backend
- 5:31database, let's jump in to start
- 5:33designing the agent system here. Okay.
- 5:36The first thing we're going to do is
- 5:37that we need to think about when will a
- 5:41user start to interact with our system.
- 5:43Right? So in this case, our user
- 5:45channels are going to be website chat
- 5:47and emails. Okay? So if you think about
- 5:50this more holistically, we're not going
- 5:52to directly allow anybody to start using
- 5:54the system. So we're going to start um
- 5:56doing some authentication first. So here
- 5:58we're going to introduce a gateway. A
- 6:01gateway is going to deal with the
- 6:03authentication single sign on or PII
- 6:06which is for privacy of the user or
- 6:09you're going to think through the rate
- 6:10limit like how many times people can
- 6:12actually use this product request for
- 6:14this product and checking for
- 6:16duplication. So all these kind of things
- 6:18are sort of we need to double check with
- 6:19a gateway. So let's add a bit of arrow
- 6:22to confirm this relationship before we
- 6:25continue to process any information. We
- 6:27need to finish authentication first
- 6:30and also in order to load the previous
- 6:32chat and emails we need to query our
- 6:34data from a relational database.
- 6:36Remember we talked about the rectangle
- 6:38data tables. So what happens here is
- 6:39basically user of fetch the chats and
- 6:44emails. Okay, we're going to start think
- 6:46about how would the agent help with u
- 6:49talking to these conversations from
- 6:51customers. Okay, imagine a customer has
- 6:54just asked a question. ask a question.
- 6:58They said, "I want to return my product
- 7:03X." You probably would think, "Hey, um,
- 7:06okay, let me try to check my policy. Let
- 7:08me try to understand uh, what do I do
- 7:10next?" Right? So, exactly for AI agents,
- 7:13the first thing we do is that we're
- 7:15going to introduce this thing called a
- 7:17router agent. What it does is that it's
- 7:19trying to understand the intent of the
- 7:21question so that it knows what to do
- 7:22next. A router agent. There we go. So
- 7:27this router agent will try to understand
- 7:29the user intent and then decide the
- 7:32agent routing. Remember that one of our
- 7:34goals is that we we need to make sure
- 7:36that the response time for 50% of the
- 7:38users is always under 1 second. The way
- 7:41we do that is that we need to introduce
- 7:44another agent called a Q&A agent. What
- 7:46this Q&A agent does is that after you
- 7:50get the first question, I want to return
- 7:51my product X. You want to allow yourself
- 7:54a little bit of time for the router to
- 7:56think through it, right? It depends on
- 7:58the exact situation like how many steps
- 8:00is going to happen right after the
- 8:02router agent. You want the conversation
- 8:04to flow like natural. So perhaps after
- 8:06you ask the first question, it will flow
- 8:08into the Q&A agent and then that will sp
- 8:11that will start to speak to the user all
- 8:12the time and they'll probably start to
- 8:14respond to the user um by saying
- 8:17something like, "Oh, thanks for asking
- 8:18the question. Let me just check it for
- 8:20you." and then start showing them like
- 8:22an thinking process. Uh respond to the
- 8:25user
- 8:27um real time. Okay. So here for the
- 8:31router agent, let's think about the edge
- 8:32case first, right? If the router agent
- 8:35realize that this customer is very angry
- 8:37like the customer say that I I must deal
- 8:40with like or they ask a very very
- 8:41complex question and we think that our
- 8:44agent decides that it's it's way u
- 8:46beyond the confidence that I have to
- 8:49deal with myself. We need to have a
- 8:50mechanism to allow to trigger looping
- 8:52the human back into this agentic system
- 8:55so that the humans will take over so
- 8:57that our overall customer satisfaction
- 8:58ray will not be affected. Right? So I'm
- 9:01just going to add in this human in the
- 9:04loop. Oh this is so difficult to use. So
- 9:07now in this case the human will start to
- 9:10approve or disapprove any request from
- 9:13the users. If the router agent decides
- 9:15now that this is a valid request and let
- 9:17me just double check if it fits with our
- 9:19policy so that I will process the rest
- 9:21of the steps for you then we will allow
- 9:23it to move on to the next step right so
- 9:26the next step we're calling it a planner
- 9:29agent let me just move it back here
- 9:32decide to take the return route action
- 9:37so what happens now is that because
- 9:39we're the router agent understood the
- 9:41intent which is to um return the product
- 9:44then it will start to call the return
- 9:46planning agent. So the planning agent
- 9:48for return will start to process what's
- 9:50going to happen next. A very important
- 9:52concept is called functional calling. So
- 9:55what will happen is that we will prepare
- 9:57a very long list of functions that will
- 10:00be available as tools for AI agents to
- 10:02select from. And I'll show you a few
- 10:04example which is
- 10:07for instance um if the user need to
- 10:10return you need to introduce a stripe
- 10:12API that will allow them to get refund
- 10:14or start paying for another product as
- 10:16an alternative a Shopify agent which
- 10:18will allow us to u pull in the
- 10:20information about how do we what's the
- 10:22track what's the order status right now
- 10:24and what is the return merchandise
- 10:26authorization right now RMA and perhaps
- 10:30for some exchanges actions as well and
- 10:32also So things like oh where's my order
- 10:34right? So we got to pull in the API tool
- 10:37that will check some third party
- 10:38logistics 3PL to to understand where the
- 10:42product it is right now in real time and
- 10:44also for example if the user is managing
- 10:47some of their customer relationship
- 10:48information in a CRM system then they
- 10:51will be able to pull in some CRM APIs to
- 10:53start writing or fetching some deal from
- 10:55their pipeline or create and update some
- 10:57tickets. Okay, these are the tool that
- 10:59we're talking about here. So if we move
- 11:01back to where we were, this planner
- 11:04agent will decide what tool we're going
- 11:06to call. And in this case, the tools
- 11:08we're going to call are the return API,
- 11:11which is from Shopify, and the Stripe
- 11:13API, which is for processing payments.
- 11:15Right? The reason why I said there's a
- 11:18bit of a catch here is that this
- 11:19planning agent really depends on um the
- 11:23logic for finishing this task. It could
- 11:26be a deterministic workflow as well
- 11:28which means that we could just use some
- 11:29rules to process the rest of the steps
- 11:32especially when the steps are very
- 11:34fixed. For example, if we know that the
- 11:36user is just going to return the
- 11:37products, then what we do is basically
- 11:40okay checkify, right? Try to understand
- 11:43if um the product is still um a okay for
- 11:47for return or for exchange, right? And
- 11:49then trigger um stripe API to process
- 11:53all the payments related information,
- 11:54right? So actually one very good
- 11:57practice is that if we can reduce the
- 12:00amount of automation of using agents for
- 12:02cases where it doesn't really need an
- 12:04agent then your system will be more
- 12:06reliable. In this case just for the sake
- 12:09of presentation I'm going to show you
- 12:11how to think through the agentic way of
- 12:13a planner for return rather than the
- 12:15deterministic way but just a very good
- 12:17thing to keep in mind. Okay so let's
- 12:19move on. So in order for the return
- 12:21planner agent to function, it needs to
- 12:23firstly check some return policies to
- 12:25make sure that it's actually fitting
- 12:27into the scenario that we're talking
- 12:29about. Right? So in this case, we need
- 12:31to finally start
- 12:34looking into vector database which will
- 12:37include things like frequently asked
- 12:39questions, policies, all these kind of
- 12:41unstructured data that are in
- 12:42paragraphs. So that it's because it's
- 12:44not easy to search in the table. So it
- 12:47will be easier for us to search
- 12:48semantically. using vector embeddings.
- 12:50This process is also called rag
- 12:52retrieval augmented generation. Right?
- 12:56It's a very fancy word and a buzz word
- 12:58at the same time. But it's important to
- 13:00understand that it's basically just
- 13:02fetching information through a semantic
- 13:05search so that we will get more relevant
- 13:08information about this company's return
- 13:10policy which the large language model
- 13:13does not even know right because this is
- 13:15a private data and we want to feed this
- 13:17private data into our AI agent. So I'm
- 13:20also going to paste the policy in guard
- 13:23rails um services here. Right. So what
- 13:26this one does is they basically want to
- 13:28double check if it fits with the policy,
- 13:30right? So there's a bit of a decision-m
- 13:33uh mechanism here just from a diagram
- 13:35perspective say check policy
- 13:38via vector DB and then this vector DB's
- 13:42information will go through this policy
- 13:44check and we're going to say okay align
- 13:47with policy inform the agent. So this is
- 13:51a little thinking process for the
- 13:52planner agent to go through before we
- 13:56trigger any functional calls. You will
- 13:59need to define very clearly in the
- 14:01prompts for this return planner agent so
- 14:04that it knows that you need to firstly
- 14:06refer to all the checks. And if you
- 14:09realize that one agent is probably doing
- 14:10too many tasks, feel free to split it up
- 14:12into smaller ones so that it will do
- 14:14exact one thing so that it does not
- 14:16hallucinate. Here we're sort of a little
- 14:19bit oversimplifying the situation but
- 14:21just to make sure that you define your
- 14:22types and define everything very clearly
- 14:25so that your prompts your agent will
- 14:27understand what it needs to do exactly
- 14:29right so after we finish the ra check
- 14:31from the vector database our planner
- 14:33agent finally understood okay we need to
- 14:36trigger some APIs okay so the first API
- 14:38they're going to trigger
- 14:40is
- 14:42Shopify API if it fits
- 14:47the policy
- 14:50call Shopify API for return and if the
- 14:54next step is about returning the money
- 14:56then we should also call the stripe API
- 14:58right here. So after the stripe API has
- 15:00been introduced which means that we have
- 15:02finished the um money return. So the
- 15:06stripe API should return the information
- 15:09back to the return planner and say okay
- 15:13um payment
- 15:15refund done. Right
- 15:18after this the return planner agent
- 15:21should be updating all the context back
- 15:24to this Q&A agent because this Q&A agent
- 15:27is really doing the conversation with
- 15:29the original user channels which
- 15:31includes the web chat or the emails.
- 15:34Right? So imagine this Q&A agent is just
- 15:36like the central brain or the CEO of the
- 15:40company who needs to do all the
- 15:41communication with their customers and
- 15:43then the rest of these people are just
- 15:45part of the organization who's doing
- 15:47their task right just that the router
- 15:49agent is sort of at a higher level
- 15:51return planner is the one that the
- 15:53router agent decides to route to right
- 15:56and then you could also route to
- 15:57exchange planner or you know where is my
- 16:00order planner all these kind of tools
- 16:02right but after this planner has done
- 16:05its job, what it should do is that you
- 16:07should always update the information or
- 16:10the context back to the Q&A agent,
- 16:13right? So the Q&A agent is like, okay,
- 16:15so um return the latest updates,
- 16:22the latest status of return
- 16:26back to Q&A agent to talk to the user.
- 16:32So this can also include situations when
- 16:35the question or the product doesn't fit
- 16:38into the policy for return. In this
- 16:40case, we will also update the latest
- 16:44information back to the Q&A agent. So
- 16:46they will be able to talk to um the
- 16:48user, right? Respond to the user real
- 16:50time. So you can see that we currently
- 16:51have a pretty solid system right here
- 16:55for the goal we have, right? So what do
- 16:57we still need to do? Few things. Number
- 17:00one is that as an agentic system, we
- 17:04should always be thinking about
- 17:06observability
- 17:08metrics and evaluations. Okay, remember
- 17:10our goal
- 17:13is that we need to make sure 70% are
- 17:15automation, 30% are looping back to
- 17:18human, right? And then the satisfaction
- 17:20rate should be over 4.5 and the response
- 17:22time should be um this much, right? So
- 17:25what we do is that we got to track the
- 17:27relevant metrics across the entire
- 17:29system to make sure that these metrics
- 17:31are being met, these goals are being
- 17:33met. Okay. So that's the first thing.
- 17:36The second thing is that it depends on
- 17:39if um this current website or this
- 17:41current client is already dealing with
- 17:43some external like say CRM systems
- 17:45perhaps there there should be some
- 17:48automatic triggers as well for us to
- 17:49write back to the CRM. Okay. So we could
- 17:53add something like this here.
- 17:56a CRM system that will fetch, write,
- 17:59deal pipeline or create, update tickets
- 18:01for their internal team to manage um you
- 18:04know some of the processes, right? So
- 18:06I'm just going to add a simple arrow
- 18:07here.
- 18:10Okay,
- 18:12cool. So we got a pretty solid um system
- 18:16right here. So again this is a very
- 18:18simple overview on how do you think
- 18:21about the agentic system um from a
- 18:23non-technical background from a
- 18:25nontechnical perspective but also I know
- 18:27that I have introduced quite a lot of
- 18:29technical concepts here too but um if
- 18:31you're a technical you realize that
- 18:33there are a lot of things that we're
- 18:33missing out right there are a lot of
- 18:35things that we didn't mention for
- 18:36example about scalability about you know
- 18:38things that are related to like the
- 18:40requests per second seasonality all
- 18:42these server side of things I think at
- 18:44the end of the day if you're a product
- 18:45manager
- 18:46you will need to discuss with your
- 18:47engineering leader anyways to figure
- 18:49these things out. But you need to have
- 18:51like this basic concept of what we need
- 18:54in order to set up this product ready
- 18:56for production. I hope this is helpful.
- 18:58I hope uh you have learned something
- 18:59from this and let me know if you like
- 19:01this kind of format of video. I'm happy
- 19:03to make more and I'll probably also be
- 19:05making some videos to explain how do you
- 19:07actually build a system like this with
- 19:08code. So stay tuned. If you like this
- 19:10video, like and subscribe and make a
- 19:12comment down below. Thanks very much.
- 19:14Cheers. Hey everyone, this is Sean. So
- 19:15today let me show you how to design and
- 19:17build rag like a pro. Rag is basically
- 19:20retrieval augmented generation. It's one
- 19:22of the most important concept in AI
- 19:24system design or for AI agents. And a
- 19:27lot of people who watch my previous
- 19:28video about AI system design really
- 19:30asked me to dive deeper into these kind
- 19:32of concepts. So today we're really going
- 19:34to dive deeper into what exactly does
- 19:36rag work and what exactly is a vector
- 19:38database. And I'll not only show you
- 19:40something like this, which is a system
- 19:42design to think through how do you build
- 19:43or design a rag system, but also I'll
- 19:46show you some live code which I'll open
- 19:48source on GitHub. So feel free to check
- 19:50out my GitHub repo. I'll also show you
- 19:52how to deploy it to a Google Cloud
- 19:54Platform GCP so that if you have a front
- 19:57end, you can just connect it into your
- 19:58app and start using it right away. Okay,
- 20:00cool. Let's jump right into it and
- 20:01start. Um, so I've got the system right
- 20:04here, but I'm going to just walk you
- 20:05through it real quick. And if you
- 20:07already know what rag is, feel free to
- 20:09jump into the coding part. First, let's
- 20:10imagine you're a user and you have a
- 20:12user channel. You either talk to
- 20:14customer support with web chat or an
- 20:15email. And then let's imagine here we're
- 20:17using the same example as the last
- 20:18video, which is a customer support for
- 20:20an e-commerce website where you're going
- 20:22to ask questions about certain policies
- 20:24about the company. And let's imagine
- 20:26like the user has asked a question and
- 20:28that question is speaking to an AI
- 20:29agent. We have an AI agent that will
- 20:32speak to the user interface. So let's
- 20:34say the user has sent a question, ask a
- 20:38question to the UI agent and then the AI
- 20:41agent is supposed to respond with an
- 20:44answer. The AI agent might not
- 20:46understand when you can return a product
- 20:49for e-commerce site or D2C brand because
- 20:51maybe some of the products are over $200
- 20:54and you cannot just return the money to
- 20:55the user and you must check a certain
- 20:57policy or you must ask a certain manager
- 21:00to approve the return, right? So, how
- 21:02would an AI agent know if this is an
- 21:04LLM? It has no access to your private
- 21:06data about your e-commerce site. Then,
- 21:08we're going to introduce this concept
- 21:09called a vector database, which will
- 21:11store some of the policies about your um
- 21:14company uh return policy. And also, we
- 21:17will use the system that will do the
- 21:19retrieval to let the agent to retrieve
- 21:22the information from the database and
- 21:24then get informed so that the LM knows,
- 21:26oh, anything over $200, we cannot just
- 21:28return money back to you. We must talk
- 21:30to a manager. Okay, so let's break it
- 21:33down and see what exactly will happen
- 21:34here. So as I mentioned, we need a
- 21:38vector database. Okay, it's a vector
- 21:40database that will store things like
- 21:42FAQs, policies, these kind of things.
- 21:44And the way that the AI agent will speak
- 21:46to it is basically
- 21:50a rack
- 21:52which is going to check the policy and
- 21:55then the way that this vector database
- 21:57is built is very simple. Okay, so we're
- 22:00going to create a frame here. The first
- 22:02one is that you might have a bunch of
- 22:04original documents. Okay, some of these
- 22:06documents could be in like PDF or could
- 22:09be like a very long paragraph. So in
- 22:11order for this AI agent system to run
- 22:13very efficiently, we need to introduce
- 22:16this concept called text to chunks or
- 22:18anything PDF to chunks so that every
- 22:20little chunk is a piece of condensed
- 22:22information so that we're not like
- 22:24overwhelming the system. Okay. Then
- 22:26later we're going to need to turn these
- 22:28chunks into a thing called embeddings.
- 22:32So what is an embedding? An embedding is
- 22:35basically an array of numbers in a very
- 22:37large dimension. What I mean by that is
- 22:39that say okay now I'm going to input
- 22:41these words into chachbt and it's going
- 22:44to give me an answer. The way chach
- 22:46understood it is not like hey I you just
- 22:48it processes all the words. Instead it's
- 22:51processing a bunch of numbers. So every
- 22:53little word that you see here
- 22:55to chache is probably a 10,00 dimension
- 22:58vector with like 0.01 0.67 0.86 all the
- 23:03way until the 1,500 something dimension.
- 23:07Okay. So that the the machine will
- 23:09understand okay so with different
- 23:11numbers at different dimension or at
- 23:14different space it means differently.
- 23:16Okay. So in math there's a concept
- 23:18called vector similarity or embedded
- 23:21similarity. What it does is that maybe
- 23:23like the word of you know French, uh,
- 23:26Spanish, Chinese, these are all
- 23:28languages. So they're probably similar
- 23:30in higher dimensions in mathematics.
- 23:32Okay, but this is not the focus of this
- 23:34video. If you're not technical or
- 23:35technical, doesn't matter. Okay, so what
- 23:37we do is that we're going to use the
- 23:38tools that already built for us and
- 23:40we're going to use it well so that the
- 23:42system will function properly. Okay,
- 23:45this process is actually called seeding.
- 23:49I'm going to show you in the code in the
- 23:50real example very soon. Okay. So after
- 23:52we did the seeding,
- 23:54you have put your policy into these
- 23:56chunks of embeddings and then feed it
- 23:58into the vector database. What happens
- 24:01next is that because now the agent is
- 24:04doing the retrieval augmented generation
- 24:06by referring to the database in the
- 24:09vector database and just searching for
- 24:11the relevant information. We're turning
- 24:13the customer question into embeddings as
- 24:15well. And we're calculating, hey, what
- 24:17kind of chunks are actually relevant
- 24:19here so that we'll be able to align with
- 24:22the policy. Let me just draw this
- 24:25diamond here real quick.
- 24:28We're going to align with this policy
- 24:33and then we're going to inform
- 24:36sorry and then we're going to inform the
- 24:39AI agents
- 24:41so that it knows what exactly is going
- 24:43on with the policy. Okay. So this is
- 24:47already a very simple system designed
- 24:50for um rag for retrieval argument
- 24:53generation and vector database. So now
- 24:55I'm going to show you a real example of
- 24:57how do you interact with a rack system
- 24:59with an use case of a customer support
- 25:01for e-commerce DTOC brand. So I'm just
- 25:04going to type in the website. So you can
- 25:05see that this is a live link. You guys
- 25:07can try it as well. Um now we're landing
- 25:09in this chatbot. Not a big surprise. But
- 25:12now we have like four documents on the
- 25:14right hand side which are all related to
- 25:15the policy of um say for example how
- 25:18about returns, how about shipping uh the
- 25:20guide of the sizing as well as the
- 25:22support for contact. And let's just do
- 25:24it in one example, which is um for for
- 25:27the shipping policy, there's free
- 25:29shipping for anything that's over $50.
- 25:31So if I just ask a question be like, I
- 25:33bought my
- 25:36shoes for $80. Can I get it shipped for
- 25:42free and send it over?
- 25:46Could have chosen a better color, but
- 25:48this is not the focus of uh this video.
- 25:50This just front end. They say, "Yeah,
- 25:52you can get free standard shipping since
- 25:54your order is over $50. According to the
- 25:56policy, it's okay as long as it's below
- 25:58$80." So, I'm going to say, "I bought a
- 26:00pair of shoes for $30, which is very
- 26:04unlikely. Uh, can I get it shipped for
- 26:08free?"
- 26:10Ask the question, and then let's see
- 26:12what the rag would say.
- 26:15Okay, it told me, "Uh, your your thing
- 26:17is not qualified because under $30 for
- 26:19free shipping. Uh, there's a $50
- 26:20threshold." Okay, that's one example.
- 26:22Let's do another example. Let's say,
- 26:24okay, there's a return policy and says
- 26:26that every item that's over $200 require
- 26:28manual approval for returns and you need
- 26:30to email the company. So, I say, okay,
- 26:32can I return the shoes that I bought for
- 26:38um $1,000.
- 26:44Now, it's doing the retrieval, doing the
- 26:45thinking, doing the similarity search,
- 26:47and it's going to tell me very soon.
- 26:49Okay. Yes, you can return the shoes, but
- 26:50since the purchase amount over $200,
- 26:53you'll need the manual approval first.
- 26:54Okay, so you can see that this entire
- 26:56conversation with this chatbot has been
- 26:58strictly following the policies on the
- 26:59right hand side. And you might argue
- 27:01that why don't I just like feed these
- 27:02information into chatbt? Why do I need a
- 27:04rag here? Well, you're right. In this
- 27:07case, we don't really need a rag. But
- 27:09imagine this policy is like instead of
- 27:11three paragraph for each, it could be
- 27:13300 pages. Imagine you're dealing with
- 27:15like legal documents or stuff like that.
- 27:17And then that would be a different
- 27:18story. Okay, so in this example, we're
- 27:20just showing you an MVP of how it works
- 27:22and how you actually build it and make
- 27:24sure you understand the concepts. And
- 27:26then if you're like scaling things up,
- 27:28there are more things that you need to
- 27:29deal with for the back end. Okay, so
- 27:32let's try another thing real quick.
- 27:34Let's say, okay, if anything is over um
- 27:38$2,000, then that would need an annual
- 27:40approval. Okay, if I save this
- 27:45and I've asked the question again, I
- 27:46say, "Okay, can I return the shoes that
- 27:49I bought for $1,000?"
- 27:57It told me, "Yes, you can return for
- 27:59$1,000 as long as they meet these
- 28:01criteria. They're unworn, blah blah
- 28:03blah. And since your purchase is under
- 28:05$2,000, you don't need any special
- 28:07approval for return." Look at this. It's
- 28:08like immediate once I like update my my
- 28:11vector database for these
- 28:12documentations. Immediately my app will
- 28:15know what exactly is a policy. How easy
- 28:18is that? Right? Imagine you have like
- 28:19300 pages of policy documents as long as
- 28:21you edit it real quick and then database
- 28:23sort of re-mbbed um the entire policy
- 28:25and your robot knows exactly how to
- 28:27answer your questions. Okay, I think
- 28:29this is very convenient and now let me
- 28:32show you how to actually build it and
- 28:33use it. Okay, you don't actually need to
- 28:35build anything. I already prepared the
- 28:37GitHub repo for you. It's open source,
- 28:39completely free. Just try it. Okay, so
- 28:41my GitHub is right here. It's got
- 28:43shenanT-ra.
- 28:45YT stands for YouTube. Rack stands for
- 28:47retrieval augmented generation. Okay, so
- 28:50if this is the first time you're going
- 28:52to touch code, don't freak out. What you
- 28:54can do is that you can just sort of you
- 28:56can honestly you can just like command A
- 28:58and copy and then paste the whole thing
- 29:00into chat GPT or into cursor and ask it
- 29:03to tell you what you're going to do. All
- 29:05right. Well, in this case, I'm going to
- 29:06show you step by step how to do this and
- 29:08launch this as a fullstack product and
- 29:10um so that you can like link it to your
- 29:12own project. You can even try it in your
- 29:14own real business. Okay. And if you're a
- 29:16business in e-commerce brands, feel free
- 29:18to let me know. Happy to do some um
- 29:20extra sessions with you uh if you're
- 29:22interested in me helping you out with
- 29:24these AI automation. Okay. Let's jump
- 29:26right into it. Cool. So, uh what we're
- 29:29going to do is that firstly I'm going to
- 29:31clone the code. I'm going to say copy
- 29:33this clone. Okay. And then I'm going to
- 29:36turn on my favorite app, which is
- 29:38cursor. Uh let me create a new window
- 29:41here. Okay. So, uh I'm going to go into
- 29:45my um
- 29:49uh so I'm going to open my project uh
- 29:52desktop developer
- 29:54and YT rag and I'm going to say okay
- 29:58YT-R
- 30:01showcase. Okay, create that. Open.
- 30:06All right, I just created this new
- 30:07folder in cursor called yt-rag-
- 30:10showcase. What you got to do is that you
- 30:12got to turn on this terminal. Okay, and
- 30:15then what you got to do is you're going
- 30:16to say get clone and then paste this in.
- 30:20All right, you see now you got this, you
- 30:22got this open source project. Okay, so
- 30:25how do we deploy this and how do we use
- 30:27this? All right, let me try to guide you
- 30:30through the readme documents. So
- 30:32firstly, we have um this architecture
- 30:34that has an app that has the core
- 30:36service, a core folder that will
- 30:39basically define the configuration and
- 30:40infrastructure. We have the models,
- 30:42pyenic models that basically defines
- 30:44what are the data types that are going
- 30:45to be required for every model or every
- 30:47API. And we're going to have a service
- 30:49model, sorry. Then we're going to have a
- 30:51service folder that's going to deal with
- 30:52all the business logic regarding rags,
- 30:55embeddings, AI agents, all these kind of
- 30:56stuff. All right. And then we have a
- 30:58main.py, which is a fast API application
- 31:00for the back end. If this is your first
- 31:02time to deal with backend, trust me,
- 31:05don't freak out. This is super easy. I'm
- 31:07not from a computer science background.
- 31:09I learned all of this by myself through
- 31:11AI. And you can do this, too. And I'm
- 31:12going to show you how to do it. Okay,
- 31:15cool. So, we already did the first step,
- 31:17which is get clone. And now all we got
- 31:19to do is got we got to go back to let me
- 31:22just set it set this up. Now we got to
- 31:24do is that we got to go to cdt-ra.
- 31:28Okay, as this said and then we're going
- 31:30to create this virtual environment.
- 31:31Literally just copy this. Okay, paste it
- 31:35in your terminal. All right, so what
- 31:37this did is that created this virtual
- 31:39environment called uh vm_yt-ra
- 31:44and then you can basically install your
- 31:45Python packages in it. Okay, so we
- 31:48already have our Python packages ready
- 31:49which is in requirements.tsx.
- 31:52We only using these few. All right, and
- 31:54just paste this in
- 32:00and you're going to install the packages
- 32:01you need for running this project. Okay.
- 32:06And then we're going to need to set up
- 32:08our vector database. In this case, we're
- 32:10using Superbase. For the vector
- 32:11database, we're going to use the PG
- 32:12vector. I know there are a lot of
- 32:14options on the market. There's Pine
- 32:15Corn, there's Reviet, and we're not
- 32:17talking about which one is better here.
- 32:19I'm just talking about how it works.
- 32:21Okay, feel free to try other tools if
- 32:23you want to. Okay, cool. Let's get
- 32:25started. Um, so the reason why we need
- 32:28to set up Superase, one is because we
- 32:30need to set up this place where it's
- 32:32going to store these policies for
- 32:34documentations. Number two is that we
- 32:36need these keys from Superbase so that
- 32:38our app is connected to the database so
- 32:40that they're communicating with each
- 32:41other. Okay, just follow me. Like trust
- 32:43trust me like this is pretty easy. We
- 32:45can do this. Cool. So let me start uh
- 32:48creating a new superbase project for us.
- 32:51Uh
- 32:53superbase.com.
- 32:55Okay. If I go to dashboard, let's make
- 32:58it bigger. I can go to YouTube. You can
- 33:01see that I already have a yt-ra. I'm
- 33:03going to create a new one just for this
- 33:05showcase. You can click on new project.
- 33:08And then I'm going to say yt-rag-
- 33:12um showcase. And then I just input in my
- 33:15database password. And let's just create
- 33:18the project.
- 33:20I'm going to delete this later. So don't
- 33:23even try to use things here because I'm
- 33:25going to show some keys here. And I'm
- 33:27going to just delete it. Don't try to
- 33:28use it. And you should set it up
- 33:29yourself. Okay. Um cool. So now we set
- 33:33up this superb basease. Let's see what's
- 33:35what's going to happen next. Okay. Let's
- 33:36come back to the GitHub. You see that?
- 33:38Wait for the project to be ready and
- 33:39then go to settings and get all these
- 33:41APIs and copy them into our um uh into
- 33:45our code. All right. So, let's I'm going
- 33:46to do it. So, go to superbase,
- 33:50come over here. I think if I just scroll
- 33:52down. Yeah, but if you just scroll down,
- 33:54you can see that there's project URL and
- 33:56API keys here. And then we're just going
- 33:58to literally going to um firstly, we can
- 34:01see that there is a file called uh
- 34:04M.ample. All right, we're just going to
- 34:07copy this and paste this again. And
- 34:09we're just going to call it M instead of
- 34:11M example.
- 34:13So that this is for production. And then
- 34:15we're going to copy the project URL
- 34:17here. Click on copy. You can see there's
- 34:19a project URL. Replace this with the
- 34:21real one. And then you can see that
- 34:23there's this Anom public key. Copy that.
- 34:26Come here. Replace the second one with
- 34:28it. Hit command S to save it. Third one
- 34:31is a service ro key. Service ro key is
- 34:33right here. Let's see. Uh let's go to um
- 34:37project setting and then we can find API
- 34:39keys and then we can see this service
- 34:41row key. All right. Going to reveal it.
- 34:43Copy it. Come back here. Paste it in.
- 34:47Command S to save it. Okay. I'm going to
- 34:49delete all these projects. So if we
- 34:50click back into table editor, you can
- 34:52see that we don't have tables now. We're
- 34:54going to set it up real quick. All
- 34:55right. One small thing is that you can
- 34:57see that in this environment variables
- 34:58right now, um there's an AI provider.
- 35:01I'm choosing anthropic. You can also
- 35:03replace it with open AAI, right? And
- 35:04then there's also like this OpenAI key,
- 35:07OpenAI embedding model. We're actually
- 35:08using this embedding model to embed
- 35:10these words. Remember we say we're going
- 35:12to turn every word into a 1,500
- 35:15dimension of numbers, right? And then
- 35:18we're going to use OpenAI. If you're if
- 35:19you're using OpenAI, you're going to use
- 35:21GBT40. If you're using Enthropic, you're
- 35:24going to use enthropic chat model, which
- 35:26is in our case, I'm going to use 3.5.
- 35:28Doesn't matter. Choose whatever you
- 35:29want. Okay. Last but not least, the
- 35:31environments development and log info is
- 35:33info right now. Okay. So, what's missing
- 35:36in the environment variable is just I
- 35:38need my open AI key and my uh anthropic
- 35:41API key. Okay.
- 35:43Um, so I'm going to just going to paste
- 35:44my own keys here. Here I'm going to
- 35:46block it. And if you don't know where to
- 35:48find it, you can just find you can just
- 35:49literally search OpenAI API key on
- 35:52Google and then you can just create an
- 35:55OpenI key. Login if you haven't logged
- 35:58in. Just click on create new keys and
- 36:00type a name here and then copy it,
- 36:02right? And then you're going to start
- 36:03using it. Enthropic is the same thing.
- 36:05Now we're going to initialize the
- 36:07database. The way we do that is we have
- 36:10a file called SQL/init
- 36:14superbase.sql. Just do command A,
- 36:16command C. Come to your superbase. Come
- 36:19to the left hand side. There's a SQL
- 36:21editor. Open it. Click in. Command V to
- 36:24paste it. Command enter to run it.
- 36:29Okay, you see, oh, your superbase
- 36:30database is ready for rag. Okay, what
- 36:32did we do exactly? I'll ver briefly
- 36:34explain. Go to the left side bar. Let me
- 36:36just make it bigger for you. Go to the
- 36:38left side bar. Click on table editor.
- 36:40You can see there's a table here called
- 36:41rag chunks. Okay, remember earlier we
- 36:45said we need to do this thing called
- 36:48seeding, which we're going to turn
- 36:49original documents into text chunks and
- 36:52then embed the text chunks. Okay, here
- 36:54I'm going to show you how to do that
- 36:55exactly. So for now you can see this
- 36:58table that has ID chunk ID source where
- 37:01does the chunk come from text what
- 37:03exactly is in that chunk what is the
- 37:05embedding for this text right when is it
- 37:08created that's all u we're just keeping
- 37:10things simple here if you go to the
- 37:12sidebar and click on database you can
- 37:14see that there's one table here listed
- 37:16and then you can also click into
- 37:18functions okay why we're in the
- 37:20functions that's because if we go back
- 37:23to uh this we actually defined some
- 37:26functions here. One function is called
- 37:28match chunks. You can just copy this and
- 37:31come back here and then search match
- 37:32chunks. You see this is basically doing
- 37:35like matching like doing the retrieval,
- 37:37right? Comparing if your query embedding
- 37:39is similar to the embedding that you
- 37:41added to the chunk. All right. And
- 37:42there's several other functions that we
- 37:44defined. Um there's also a function
- 37:46called
- 37:50There's also a function called get chunk
- 37:52stats, right? Let's just search this row
- 37:54here. Here real quick. Yeah. And you can
- 37:56also click on these three dots and click
- 37:58on edit function to edit things. And if
- 37:59you're familiar with the SQL, you will
- 38:01know what's going on. It's going to it's
- 38:02going to s select from the rack chunks
- 38:04and count how many total chunks you
- 38:06have. Uh what are some unique sources
- 38:08you have and what are some of what is
- 38:10the total maximum sorry and what is the
- 38:12latest time when you update the chunk
- 38:13table. Okay. So this is called superbase
- 38:15functions. Again, this is not the focus
- 38:17of the video. If you're interested in
- 38:18superbase, feel free to watch my
- 38:19previous videos on superbase. Let's move
- 38:21on. Continue.
- 38:23Come back to this. You can see that we
- 38:25initialized the database. Um so here one
- 38:29thing to interesting to explain is that
- 38:31we're using this thing called pg vector
- 38:32which is part of postgress and it's
- 38:34basically a way to save vector database
- 38:36right again feel free to use pine con we
- 38:39or lchain stuff all these things work
- 38:41right doesn't matter you don't have to
- 38:42use this one okay so with the rack
- 38:44chunks crael currently we're using um a
- 38:47vector of dimension 372 dimensions so
- 38:50that's like doubling of 1,500 when I
- 38:52said it okay so uh now let's move on
- 38:57uh oh sorry now let's move on so you can
- 39:00see that I have this thing called test
- 39:02setup.py okay let's come back here and
- 39:05see where is it we have this file called
- 39:07test setup.py Pi. So what this test up
- 39:09do does is that I'm doing a few things.
- 39:11Firstly, I'm going to import the
- 39:12modules. Second, I'm going to check the
- 39:14configuration and then I'm going to
- 39:16start doing the database connection and
- 39:18then I'm going to start doing the schema
- 39:19validation. And now I'll do the seeding
- 39:22documentation. Look at this. I'm doing
- 39:23the seeding documentation so that we are
- 39:25seating the documents into our database
- 39:28so that later we can do the rack query.
- 39:30I'm going to show you step by step.
- 39:31Okay. Uh feel free to just run this doc
- 39:33if you're familiar with this. But if
- 39:34you're not familiar with this, I'll show
- 39:35you what exactly we're going to do.
- 39:37Okay, so firstly, how does seating work?
- 39:40Okay, so the main.py is an app
- 39:42slashmain.py.
- 39:45Okay, so this is basically a backend
- 39:47endpoint. What it does is that sorry,
- 39:50maybe it looks scary if you're not
- 39:51technical, but I'll explain everything
- 39:52as I as I as I always mention. Okay,
- 39:55first you have this live span
- 39:57definition. What it does is basically
- 39:58just going to try to, you know, connect
- 40:00to the database, initialize the schema,
- 40:02all these kind of stuff. And then we
- 40:04have this thing called fast API. We're
- 40:05defining fast API as this name called
- 40:08app. So anything with app app do
- 40:10something app do something that's fast
- 40:12API. Okay. And we're going to add this
- 40:14middleware which is basically saying hey
- 40:16local host 3000 local host 30001 they're
- 40:18accessible to this back end. Because if
- 40:20you're building like the front end with
- 40:21an XJS you'll be able to you know use
- 40:23localhost 3000 or recell as your URL.
- 40:27We're just basically telling the back
- 40:29end that these URLs are access can
- 40:31access you right so that there's no
- 40:33access issues.
- 40:36Um and then we have a chat interface uh
- 40:39which is an HTML. Here we're not using
- 40:41X.js. We're just using a plain front
- 40:43end. Okay. And then um what's important
- 40:46is that we can just sort of like start
- 40:48trying things. Okay. So what I'm going
- 40:50to do is that remember to do source
- 40:53vimt_ra
- 40:56bin slash uh activivate. So, what you
- 41:00can do is just you can type in uvorn
- 41:03main app-reload.
- 41:06After we run this, we're going to get uh
- 41:08a live backend called 127.0.0.18000.
- 41:16Okay. So, if I just do commandclick on
- 41:18this thing, it's going to show me this
- 41:20thing called welcome to Rack AI Asian
- 41:22backend version blah blah blah blah blah
- 41:23blah blah. Okay. And a quick way to
- 41:26check it is that there's this thing
- 41:28called um basically you can check either
- 41:30the health of it or you can do something
- 41:32fun like for me I did like slashgreet
- 41:34slashname so that it will tell me uh if
- 41:37it actually worked. So if you go to this
- 41:39URL and do slash greetan
- 41:42it's going to tell you hey Sean I think
- 41:44you're great. Okay and if I just greet
- 41:46like Donald I think you're great. So so
- 41:49this is like working. Okay, this is
- 41:51working. And uh or you can just type in
- 41:54uh slash health,
- 41:58right? And then it's going to tell you,
- 41:59oh database connected is true. All
- 42:01right, so now let's move on.
- 42:05So here are the three things that are
- 42:06very important. The first one is an
- 42:08endpoint called the slash documents.
- 42:10This is where we're going to save all
- 42:11the documents. Okay, so if you go to
- 42:13slash documents, it's going to return to
- 42:15you what documents you have. All right,
- 42:17so let's try that. So if you do slash
- 42:21documents you can see all the documents
- 42:23we have I'll show you in the code where
- 42:26is it where is it so if I just do double
- 42:29click command default command c command
- 42:31f command v search you can see that we
- 42:34imported from this okay which is where
- 42:37is it
- 42:39uh from data slashdefault documents so
- 42:42from data folder where's data folder
- 42:45there we go data folder/default
- 42:47documents So these are the documents for
- 42:49the policies that we're going to input.
- 42:51All right. So we're going to have a
- 42:53policy for return as I if you still
- 42:56remember in our app we have a policy for
- 42:58return v1 policy for shipping sizing
- 43:01guide and support contact. Okay. So we
- 43:04have basically the exact same thing.
- 43:05Policy return policy for shipping policy
- 43:09for sizing and policy for support
- 43:11contact. Okay. So if we go back to
- 43:13main.py.
- 43:16So if you go to go back to app/main.py,
- 43:19if you call the API of u the URL/d
- 43:23documents, it's basically going to show
- 43:25you the default documentations as I
- 43:26showed you earlier. Okay. Another
- 43:28endpoint is called slash seed. What this
- 43:30one does is they're going to take the
- 43:31input of the documentation and then
- 43:33start showing you, you know, hey, we're
- 43:35going to we're going to we're going to
- 43:36like turn your documents into chunk ID,
- 43:40source, and text. Remember, this is how
- 43:42we're going to store it in the database.
- 43:43All right. So this API I'm literally
- 43:45going to show you how to turn these
- 43:46documentation into the database. All
- 43:48right. So what it does is that it
- 43:50firstly split your thing into chunks.
- 43:54All right, split thing into chunks and
- 43:56then it's going to show you and then
- 43:59it's going to run this function called
- 44:00rack service C documents. You can do
- 44:03commandclick into this. Sorry.
- 44:07And then it's going to run this function
- 44:09called the rack seed documents. Okay.
- 44:12And then it will get to this the
- 44:14inserted account. So, so how did how did
- 44:16the whole thing happen? Right? Let me
- 44:18show you. So, if you do if you run this
- 44:21API, if you run this API, what it does
- 44:24is that it's going to run this function
- 44:25for you. Okay, let's find out where is
- 44:27this function. Command C, command F,
- 44:29command V. All right, we define it in
- 44:33services.rag.
- 44:35Right, we have a folder called services.
- 44:37This is basically where we're going to
- 44:38save all these AI agents, right? Or
- 44:40agent related stuff. So let's go go to
- 44:42the rag.py. Within the rag.py, remember
- 44:46like we I mention remember I mentioned
- 44:48that we have a function called seed
- 44:49documents. What this one does is that
- 44:51it's going to turn your documents into
- 44:53chunks and turn chunks into embeddings.
- 44:56Okay, so let's briefly go through it. We
- 44:58have a function here that say okay we're
- 45:01going to use the chunk function to chunk
- 45:03the documentations and then we're going
- 45:05to embed the documentation using
- 45:07embedding service which is imported from
- 45:09embedding.py. Pi. And last but not
- 45:12least, we're going to combine the chunks
- 45:13into embeddings so that we can insert it
- 45:15into the database. Okay. Um, let's check
- 45:18the embedding. Let's check the
- 45:20embedding.py real quick, too. In the
- 45:22embedding, what we do is that you can
- 45:24see that here we have embed text. What
- 45:27it does is it uses this OpenAI client
- 45:29with embeddings to create the embeddings
- 45:31based on the text input. Remember, we
- 45:33already cut the text into smaller
- 45:35chunks. So, it's going to it's going to
- 45:36embed every text into embeddings. Okay.
- 45:39I know I say a lot of tongue twisters.
- 45:41Let's just run it and see what happens.
- 45:43All right. So all you got to do is that
- 45:45if you have defined properly of your
- 45:48default documentations, what we can do
- 45:50is that we can just do test the setup to
- 45:53set up the whole thing. Okay. So come to
- 45:55your terminal again if you're are not in
- 45:58your virtual environment do source
- 46:00vt-_rag
- 46:04bin activate. Okay. And then we're going
- 46:06to do a Python test setup. Now it's
- 46:11going to run this setup and I'll show
- 46:13you what exactly we're doing.
- 46:16So we imported the modules testing the
- 46:19configuration did the database
- 46:20connection schema documentation seating.
- 46:23You can see that it successfully seated
- 46:25the four document chunks and then it
- 46:27tested the rag. It worked. Okay. And it
- 46:29passed everything. Okay. So now if we go
- 46:34back to the database,
- 46:36go to the left sidebar, table editor,
- 46:40you're going to see all four
- 46:41documentations are here. All right.
- 46:44Policy, return, shipping, guide,
- 46:46contact, and you can see where the
- 46:48sources from and you can see what the
- 46:50text exactly is if you double click on
- 46:52it. Okay. Anything that items over $200
- 46:54require manual approval for return. And
- 46:56then there's this embedding that is
- 46:583,000 columns of dimension. And if you
- 47:01do this slash chat, you were to turn it
- 47:04on. And now you have these policies
- 47:06right here. Okay. So this is literally
- 47:08using what we defined here. Okay. So let
- 47:12me prove it to you by making this a
- 47:14little bit smaller. Making this a little
- 47:16bit smaller.
- 47:18And our eye should be focused on our
- 47:20terminal here. Okay. So what we got to
- 47:22do is that I'm just going to say okay
- 47:25what is what is my return policy and hit
- 47:29send. You can see that here it's running
- 47:32because now it's doing the vector
- 47:34database retrieval and you can see that
- 47:36it's using anthropic. All right. And
- 47:39then it did the query processing with
- 47:41one citation. So that means it find one
- 47:44reference one citation. All right. Which
- 47:46is the policy that we refer to return
- 47:49policy. Okay. And I just asked the
- 47:51question again. Can I return for return
- 47:54my product that was $800?
- 48:00You see now it's doing this another uh
- 48:02retrieval again. And then you see they
- 48:03it tells me since your item is already
- 48:05$200, you'll need manual approval for
- 48:07the return. Okay. So now we prove that
- 48:09this rack system already works, right?
- 48:11It already works.
- 48:14One more thing I want to show you is the
- 48:15chat. If you click on serviceshat.py,
- 48:18Pi. You can see that we're currently
- 48:19doing this function called generate
- 48:21answer. What this one does is that
- 48:23there's a system prompt, right? It says
- 48:25that you're a helpful AI assistant for
- 48:27customer support. And there's some
- 48:28important rules. You basically say you
- 48:30can only answer questions for policy,
- 48:32return, shipping, and sizing. And you
- 48:34should for questions outside of
- 48:35knowledge base, you can just politely
- 48:37reject that. All right. And then we can
- 48:39just test it real quick too. You can say
- 48:42uh who is the US president now?
- 48:52Yeah, I apologize. It's not relevant.
- 48:55So, he's not going to answer me. Okay.
- 48:57Feel free to play around with this
- 48:58GitHub if you're technical. And if
- 49:01you're not technical, if you follow this
- 49:02exact steps that I talked about, you can
- 49:04set this up for your business. Happy to
- 49:06help you to set up for you if you're
- 49:07interested in this kind of AI
- 49:09automation. And the last step I do is
- 49:12I'm going to make this live. I always
- 49:14always want to make product live because
- 49:15I feel like that's the final step you
- 49:17do. Like that's that's the core of
- 49:19building a product. It needs to be live.
- 49:21It needs to be able to be used by
- 49:22someone. Let's do that. Cool. So, um
- 49:27what I'm going to do is I'm going to go
- 49:29to my u Google cloud. So, I normally use
- 49:32Google Cloud for this. And I'm I already
- 49:35have a project. If you don't have a
- 49:36project, feel free to click on here and
- 49:38then click on create a new project.
- 49:40Okay? But that's not in the scope of
- 49:41ours. So, I already have a project. I'm
- 49:43going to go right into Cloud Run on the
- 49:45left hand sidebar. Okay. And then I'm
- 49:47going to select a project which is the
- 49:49project I created. And you can see all
- 49:51the project I have here. And I already
- 49:53have one called YT-Rack. And the rest of
- 49:55them are my previous projects. So what I
- 49:57do is that I first need to commit the
- 49:59code to the GitHub. And then we can
- 50:01connect it through GitHub into Google
- 50:04Cloud. Okay. So let's try this real
- 50:07quick. Let's come back to cursor.
- 50:10And one thing I didn't show you earlier
- 50:12is how to connect to GitHub. And if you
- 50:14are creating your own thing, you can
- 50:16also come back to your folder and do
- 50:18remove- rf uh git. Okay. Hit run. Create
- 50:22a new repository. Okay. And I'm just
- 50:25going to say yt- um rack slash um
- 50:31private
- 50:32showcase. Okay. Create the repository.
- 50:37Get remote ad. Make sure you make sure
- 50:40you navigate your folder into YT-R and
- 50:44then get branch man.
- 50:47Okay, let's come back here. Command R to
- 50:49refresh it. Cool. This is good. Um, so
- 50:52now we got this. All right, we got the
- 50:55whole thing. Let's come back to Google
- 50:57Cloud.
- 50:58Let's click on connect repo in Cloud
- 51:02Run.
- 51:04Let's set up the cloud build name. All
- 51:06right. And then I'm gonna click on here
- 51:09to repository. I can say yt-rag dash.
- 51:12You can see I have two of them but I'm
- 51:14going to use the private showcase
- 51:16understand. Next
- 51:24okay we're going to use this thing
- 51:26called docker file. Okay. What this one
- 51:28does is basically telling Google cloud
- 51:29that we need you need to install
- 51:31requirements.txt
- 51:32and the main working directory is app.
- 51:34Come back here and just hit on save.
- 51:39Okay. And I'm going to allow public
- 51:42access because it's not a big deal right
- 51:44now. And then if I come back here,
- 51:47there's an important thing called
- 51:48variable and secrets under containers.
- 51:51We need to add the variables here.
- 51:52Select the whole thing. Paste it here.
- 51:55Okay. And just hit on create.
- 52:02Now you're deploying the app. Okay. So
- 52:04this will be a URL.
- 52:07All right. So now creating the service
- 52:09by the way. Now if you go to your GitHub
- 52:12and if you refresh it,
- 52:16you can see that something's running
- 52:17here, right? Because we're connected to
- 52:19GCP through Docker, right? Can click on
- 52:21the details. You can see that it's in
- 52:24progress. So it's it's very well
- 52:25integrated. So if you click on view more
- 52:27details on Google Cloud, you can see how
- 52:29exactly it's been built. You see it's
- 52:31very well connected. If you click on
- 52:32this commit thing, it will lead you back
- 52:34to your GitHub for the exact commit. Oh,
- 52:37good. So, now we're officially live. And
- 52:39then we click on YT RA private showcase.
- 52:43You see, we have this URL here. Copy
- 52:45this. Command T, command V. Good guys,
- 52:49this is live. And then let's do slash
- 52:51chat. Cool. We got our policies here.
- 52:55So, for return, let's say it should
- 52:58become if anything is over $8,000.
- 53:01All right. If I hit on save, by the way,
- 53:04double click here again. You see it
- 53:06changed into 8,000 immediately. And they
- 53:07updated this embedding as well,
- 53:09immediately. Okay. So, this template
- 53:11really works well. Okay. So, now if I
- 53:13ask the exact same question again,
- 53:16it's doing retrieval on an updated
- 53:18database, vector database, and it's
- 53:20going to tell me uh yes, you can return
- 53:23it and purchase 8 1,000, which is lower
- 53:26than 8,000 is fine. This is the end of
- 53:28the video and I hope you enjoyed it. I
- 53:31hope this is straightforward. I hope
- 53:33this is like easy to follow and again
- 53:37like I I would like to hear any feedback
- 53:39from most of you guys and a lot of
- 53:42people ask me for making more videos
- 53:43regarding you know showing you AI agent
- 53:45system design and people also ask me
- 53:47about deployment. So I'm trying to keep
- 53:48the balance. So let me know if this
- 53:50format of video works for you. like I'll
- 53:52show the system design at the beginning
- 53:54and I'll show you like the exact code
- 53:56example and open source the code on
- 53:58GitHub and show you how to deploy it.
- 54:00Let me know if this kind of format of
- 54:01content is really helpful for you. I
- 54:03don't think anyone else I don't I didn't
- 54:05find any other YouTubers or influencers
- 54:07doing this. Um so I'm sort of also
- 54:10experimenting this myself and at the
- 54:13same time if you're a business if you're
- 54:15are running an e-commerce website or D2C
- 54:18brand um I'm running my own startup
- 54:20called automatis.io IO which is a which
- 54:23is basically an AI agent system that
- 54:25helps people selling physical products
- 54:27to manage their sales leads. And I'm
- 54:29also open for discussing any AI
- 54:31automation requests or demands from a
- 54:33business perspective. Just let me know
- 54:35and I have all my social media contacts
- 54:37down below in the description and in the
- 54:39comment section. So feel free to reach
- 54:41out. Hey everyone, this is Sean. Today I
- 54:43want to talk about how to build a strong
- 54:44AI agent system.
About this transcript
This page contains the full transcript of 8 сентября 2026 г. by Andrei, generated from the public captions YouTube serves with the video. The transcript has 10,044 words across 1,427 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.