Frontend System Design Explained w/ Senior Engineer (Microfrontends, Monorepo, MCP UI, Reactjs) — Transcript
Full transcript
- 0:00As AI is getting better at writing
- 0:02front-end code, your only choice to stay
- 0:04relevant and competitive as a front-end
- 0:06engineer is to level up your system
- 0:07design and architecture skills. The
- 0:09problem is that most system design
- 0:11content is way too focused on the
- 0:13back-end while completely ignoring the
- 0:15front-end. So today I'm going to break
- 0:17down every front-end system design
- 0:18concept you need to know from
- 0:19micro-front-ends architectures to system
- 0:22design patterns like back-end for
- 0:23front-end and then go into mono-repos
- 0:26and rendering strategies like
- 0:27server-side rendering. I'll also show
- 0:29you what is the best way to use AI and
- 0:31agentic coding to deploy these systems
- 0:33to production. Now let's start with our
- 0:35first front-end system design concept,
- 0:37micro-front-ends. Now the traditional
- 0:38front-end application will start as a
- 0:40front-end monolith. That means pretty
- 0:42much everything is in the same place.
- 0:44And when it comes to micro-front-ends,
- 0:46it basically means that we will split
- 0:48this application into independent
- 0:50front-end applications and then put them
- 0:52all together. But before we even talk
- 0:54about micro-front-ends, we need to
- 0:56understand microservices and that is a
- 0:58back-end architecture style. And so
- 1:00looking at our client-server model, both
- 1:02the server and the client will be
- 1:04monoliths. That is single applications.
- 1:06And at that point we keep adding
- 1:08features to it because you're in the
- 1:09initial phase of the project. But as you
- 1:12scale, more and more people will have to
- 1:14work on the same code base and your
- 1:16development team becomes huge. And I've
- 1:18worked with developers that told me
- 1:20their development team was 40 people and
- 1:23their daily stand-up took 1 hour and a
- 1:25half just for everybody to speak for 1
- 1:27minute and a half. And that's just not
- 1:29sustainable. So that's where
- 1:31microservices and micro-front-ends come
- 1:33into place because this kind of
- 1:34architecture doesn't scale from a
- 1:36development point of view. Now you've
- 1:37probably heard about the term two pizzas
- 1:40team where Jeff Bezos said that team
- 1:42shouldn't be bigger than five to nine
- 1:44people, which is as many as you can feed
- 1:46with one pizza. Now the two pizzas team
- 1:49becomes a one pizza team with AI.
- 1:52Because nowadays you can do the same
- 1:53work with a lot less people using
- 1:56agentic coding. And so, the ideal size
- 1:58of the team becomes smaller and smaller.
- 2:00Now, do remember that even if coding
- 2:02agents like Cloud Code drive
- 2:04implementation time to zero, they do
- 2:06increase the verification time. So,
- 2:08being an AI savvy front-end engineers it
- 2:10means you design system that decrease
- 2:12the verification time and minimize
- 2:14architectural drift while they maximize
- 2:15velocity. And so, it's not anymore about
- 2:18implementing fast, but rather building a
- 2:20system that is extremely easy to verify.
- 2:23And we'll see how that changes
- 2:24architecture later on. But, keep this
- 2:26principle in mind. And so, in the ideal
- 2:27case we are able to split our monolithic
- 2:29team into feature teams that own a
- 2:32specific feature. And those teams are
- 2:35full stack and they work with coding
- 2:37agents, but they can deliver
- 2:38independently and that's how we can keep
- 2:40our product moving forward. And so,
- 2:42we'll have these different feature
- 2:43teams, but it's also shared
- 2:45infrastructure and that's owned by the
- 2:46platform team. You might have front-end
- 2:49tooling or pure infrastructures like the
- 2:51servers. So, you might have a team for
- 2:53example that is building a design system
- 2:55that the feature teams are using. And
- 2:57all those teams are smaller and they all
- 2:59work with coding agents. That is the new
- 3:01norm in development teams. Now,
- 3:03according to Conway's law, the software
- 3:05architecture of a system will look like
- 3:06the organization of the people. So, if
- 3:08what we want is smaller independent
- 3:10teams, then we need to somehow break our
- 3:13system into small independent vertical
- 3:17slices. And this truly starts in the
- 3:19back-end where we take our monolith and
- 3:21we split out the independent modules to
- 3:24build microservices. Those are
- 3:26independent small back-ends, micro
- 3:28back-ends that expose their own API and
- 3:31can be deployed independently. The code
- 3:33base is independent and you have
- 3:34independent teams that don't need to
- 3:35talk to each other extending those and
- 3:37they only communicate between each other
- 3:39through APIs. And so, a typical
- 3:41microservice blueprint will look like
- 3:43something like this. We have an API,
- 3:44there's a business logic layer, then
- 3:46there's a persistence layer and usually
- 3:47a database. In this case I added a
- 3:49Postgres DB, but it could be a no sequel
- 3:51DB. And so, this is the building unit of
- 3:53your application, and big applications
- 3:55will have thousands of those in
- 3:57productions. A couple of years ago, I
- 3:59used to work for this financial company,
- 4:00and we had around 1,000 microservices in
- 4:03production, and you had different teams
- 4:04only different microservices. My team
- 4:06owned 13 of them. And by the way, if you
- 4:08want to see where do you stand across
- 4:10the full stack, there's a free
- 4:11assessment that you can take in the link
- 4:12below, and you basically understand how
- 4:15much of this full stack knowledge you
- 4:17actually have, where is the gap, and we
- 4:19added a new AI section because a lot of
- 4:21companies are starting to ask AI
- 4:22questions in front-end engineering
- 4:25interviews, also. So, take the
- 4:26assessment and you can see more or less
- 4:28where do you stand in the market and
- 4:29what gaps you need to close. Link it's
- 4:31in the comments. Now, moving on, we
- 4:33split our back end into microservices,
- 4:35but our front end is still a monolith.
- 4:37So, even if our back end teams now could
- 4:40ship independently and be much smaller
- 4:42and go much faster, the front end is
- 4:43pretty much the same. And the problem
- 4:45with the front end is that you have so
- 4:47many people pushing code to the same
- 4:49client that it's so easy for someone to
- 4:51make a mistake or bring the whole system
- 4:53down or change a global CSS rule that
- 4:57then affects everybody. So, as of now,
- 4:59our front end team is still very, very
- 5:00big and totally not sustainable. It gets
- 5:03to a point where we just cannot scale
- 5:05because there's just too much
- 5:06communication overhead between all those
- 5:07people needing to coordinate as they
- 5:09work on the same code base. And that's
- 5:11why we go back to micro front ends,
- 5:12where we have a front end monolith and
- 5:14we split it in different independent
- 5:16front end applications. Now, how do we
- 5:18put all those together? Well, basically,
- 5:20we'll have a shell. And the micro front
- 5:22end shell is the one that takes care of
- 5:24the global functionality, things like
- 5:27auth, routing, language, global state,
- 5:30you know, are the users logged in or
- 5:31not, it's all handled by that. And all
- 5:34the others model applications, they live
- 5:36inside this one. They're basically
- 5:37loaded by this micro front end shell.
- 5:40And this micro front end shell, it's
- 5:41also on passing this global state inside
- 5:44those micro front ends. But remember,
- 5:46those applications are now deployable
- 5:48independently. That means you might have
- 5:49a domain that it's
- 5:51header.theseniordev.com,
- 5:52where, for example, we would host only
- 5:54our header application, and you could
- 5:56load that independently. And then you
- 5:57have the product page and the cart page,
- 5:59but then you have the shell that puts
- 6:00them all together. So, if you look at a
- 6:02big software company like amazon.com,
- 6:04you could basically split the header,
- 6:06and then maybe take out the product page
- 6:09and the payment into two different micro
- 6:11front-ends. But, to be honest, I feel
- 6:13like they have even more granularity
- 6:15there, and they probably split their
- 6:16front-ends into even smaller and more
- 6:19focused micro front-ends. So, once you
- 6:21apply micro front-ends and microservices
- 6:23together, you end up with a micro
- 6:24front-end that will speak to one or
- 6:26several microservices in the same
- 6:28domain. And this is what we call a
- 6:30vertical slice. And this system can now
- 6:33be extended by an independent team. So,
- 6:35we can start releasing changes
- 6:36independently. Our releases are not a
- 6:39bottleneck for another team, and we can
- 6:41work against the APIs of other teams.
- 6:44This is how pretty much every big
- 6:46software company operates. And this is
- 6:48what in software architecture they call
- 6:50vertical slicing, where your features
- 6:52are vertically owned by one team. And
- 6:55I'm going to go even farther and tell
- 6:57you that if you are an engineer, and
- 7:00you're doing only front-end or back-end,
- 7:02you have to move as fast as you can
- 7:03towards becoming a vertically integrated
- 7:06engineer, basically a team of one that
- 7:08can work with a coding agent across the
- 7:10full stack. And we'll see in a second
- 7:12why that happens. But, that's the
- 7:13transition you're going to need to go
- 7:15through to survive in today's market.
- 7:17So, when it comes to traditional
- 7:18front-end engineering and front-end
- 7:19development, we see that moving towards
- 7:20two extremes, where you're either more
- 7:22of a full-stack person that works in a
- 7:24feature team, so they work across these
- 7:26vertical slices, and that's why you need
- 7:27to know a bit more about full-stack. By
- 7:29the way, we do have full-stack videos on
- 7:31this channel. If you want to know, as a
- 7:33front-end engineer, transition slowly
- 7:34towards a full-stack. Or, you're going
- 7:36to be really good at the front-end, and
- 7:38then you're going to the infra team, and
- 7:39you start building it at the shell, or
- 7:41you build a design system, and that's
- 7:42where you do spend a lot of time, you
- 7:44know, doing traditional front-end work
- 7:46with just CSS, and building the building
- 7:48blocks that the feature teams will
- 7:49implement. And pretty much both of those
- 7:52people will work with coding agents.
- 7:54Now, a quick tip about AI and AI coding,
- 7:57micro front-ends and microservices
- 7:58decrease the blast radius of a
- 8:00production bug, and they could use the
- 8:02cognitive load of code changes because
- 8:04you have smaller and localized PRs. You
- 8:07also get less context, which means your
- 8:09AI coding tool would be much more
- 8:11efficient because you work on smaller
- 8:13pieces and the risk is less. So, it's a
- 8:15really good strategy when you have a lot
- 8:16of people pushing a lot of code very
- 8:18frequently as we do with AI to move
- 8:21towards micro front-ends and
- 8:22microservices because the risk is lower
- 8:25and it decrease what we call the
- 8:27verification time. It's just easier to
- 8:29go on the quality when the surface area
- 8:31of the changes you are making is
- 8:32smaller. So, micro front-ends and
- 8:34microservices are usually the way to go
- 8:36if you're building fast with AI. Now,
- 8:37let's move to our next concept, which is
- 8:39only API gateway. Now, whenever we have
- 8:41a micro front-end and a back-end
- 8:43service, we need to end up making a lot
- 8:45of requests. And those requests have
- 8:47what we call edge functions. They need
- 8:49caching, HTTPS, authentication, content
- 8:52negotiation, rate limiting for security,
- 8:54and so all this functionality at the API
- 8:56level it's pretty repetitive. So, having
- 8:59to implement this both on the front-end
- 9:00side and on the back-end side every time
- 9:03we plug in a new microservice to the
- 9:05micro front-end ends up adding a lot of
- 9:08overhead. And it becomes a lot harder to
- 9:10do this if you have a lot of back-end
- 9:11services. So, a typical solution is to
- 9:14add what we call an API gateway, and
- 9:16that would be your gate into the
- 9:18back-end systems. And the advantage of
- 9:20an API gateway is that the client would
- 9:22only implement the edge functions once.
- 9:25So, they basically only implement the
- 9:26HTTPS handshake once, or caching, or
- 9:29rate limiting, and then once the request
- 9:31goes to the API gateway, it gets
- 9:33forwarded to the back-end services,
- 9:34which don't have to care about all this
- 9:36functionality. They can just care about
- 9:38their own logic. And all this usually
- 9:40lives in a virtual private cloud, which
- 9:42means it's totally secure because the
- 9:44only way to go into it is through the
- 9:46API gateway. The other cool feature when
- 9:48you have an API gateway is that you can
- 9:49make the communication between the
- 9:51client and API gateway HTTPS, but the
- 9:53communication between the microservices
- 9:55and the API gateway HTTP because you
- 9:57don't need that security anymore cuz
- 9:59you're in a closed environment. And the
- 10:01advantage here is performance because
- 10:03HTTPS it's likely less performant than
- 10:06the HTTP because you need more round
- 10:08trips to do the HTTPS handshake. So, in
- 10:10general, once you have a couple of
- 10:12microservices in production, it's a good
- 10:13idea to add an API gateway both for
- 10:15security and performance and also less
- 10:17complexity in the front end. Now, let's
- 10:19move on to our next pattern, very
- 10:20similar to the API gateway, we have the
- 10:22backend for frontend. Now, let's
- 10:23remember our setup from before where we
- 10:25have our client and our API gateway and
- 10:27all our microservices. The problem here
- 10:30is that the client has to make a call to
- 10:32all these different microservices that
- 10:34might have a different API and there's
- 10:36so many fetch calls. And if you're in
- 10:37the frontend team, whenever you need a
- 10:39new feature that needs even the
- 10:40slightest backend change, you need to go
- 10:42and talk to that backend team and figure
- 10:44out if that will be a priority for them,
- 10:46add it to their backlog, and maybe
- 10:48something gets done. So, it ends up
- 10:49being very, very slow. And one of the
- 10:51biggest challenge that some of the
- 10:52engineers we work with at the senior dev
- 10:55that work in bigger companies is that
- 10:56it's so hard for them to get things done
- 10:58because they have to talk to key
- 11:00different backend teams that have their
- 11:01own backlog and their own priority. So,
- 11:03they just don't want to implement that
- 11:04small API field that they need. So,
- 11:06better solution is to somehow own your
- 11:08backend as a frontend team, as a client
- 11:11team. There's also a lot of complexity
- 11:13when you have to implement all these
- 11:15different APIs because they all look
- 11:17different, you need a different client
- 11:18and SDK for all these different APIs.
- 11:20And so, that's more frontend complexity.
- 11:22So, the solution for that is to add what
- 11:23we call a backend for frontend. This is
- 11:26a backend that will take the frontend
- 11:28request and then just forward it to the
- 11:30microservices. But the advantage here is
- 11:32that whenever you have to implement a
- 11:34feature, you can act as a full-stack
- 11:37developer when you're in the front-end
- 11:38team. So, basically, you have all those
- 11:40microservices, you integrate, like to
- 11:43change the way you equals them from the
- 11:44back-end for front-end, and then you
- 11:45implement your feature, and you kind of
- 11:47own end-to-end the client-side feature.
- 11:50And the back-end teams, they can just
- 11:51work in isolation on the different
- 11:53microservices. So, the front-end team
- 11:55ends up owning both the client and the
- 11:57back-end for front-end, and back-end
- 11:59teams they develop pure back-end
- 12:01services. And this is very important for
- 12:02you as a front-end engineer because it
- 12:04means that you do need to know, at least
- 12:06at the high level, how to extend a
- 12:08microservice, how to build a back-end
- 12:09for front-end. You need to know API
- 12:11design, and you need to know a little
- 12:13bit about GraphQL, for example, which is
- 12:15a great technology to build back-end for
- 12:17front-ends. This is a requirement for
- 12:19all the front-end engineers that you see
- 12:21working at bigger companies, which are
- 12:23usually the ones that also have the best
- 12:25conditions and the most exciting work.
- 12:27So, even if you're in the front-end,
- 12:28make sure that you can also get things
- 12:30done in the back-end. Now, the advantage
- 12:33with BFFs is that you can adapt them to
- 12:35a specific client. So, if in the future
- 12:37we end up building a mobile app that has
- 12:39totally different requirements to our
- 12:41desktop app, we don't need to duplicate
- 12:43API endpoints. So, with back-end for
- 12:45front-ends, every client will have its
- 12:47own dedicated back-end, which means
- 12:49every client team can work independently
- 12:51and is completely decoupled from the
- 12:53back-end services that can just focus on
- 12:55their own service. And so, basically,
- 12:57whenever the desktop client, for
- 12:59example, goes to gateway, they get
- 13:01redirected to the desktop back-end for
- 13:03front-end. And for mobile, we have a
- 13:05completely different API, a completely
- 13:07different mobile back-end for front-end.
- 13:08It uses the same back-end services, but
- 13:10it might expose a total different API to
- 13:13consume them, just because the data
- 13:15needs in mobile are usually different
- 13:16from desktop. Where in desktop, you want
- 13:18to fetch a lot of information at once,
- 13:20but in mobiles, we have smaller screens,
- 13:21so you need different records. And if
- 13:23you try to put all those things together
- 13:25in a single API, it will become bloated,
- 13:27or it will end up not satisfying one of
- 13:29the requests. For example, you'll have a
- 13:31great API for desktop, but when it comes
- 13:33to mobile, you'll fetch too much data, a
- 13:35lot of data you don't really need
- 13:36because the interface is a bit
- 13:37different. Or if you make the smaller
- 13:39endpoints, when it comes to desktop,
- 13:40you'll have to make many requests to get
- 13:42the same data. So, you'll have
- 13:44over-fetching or under-fetching, and
- 13:46it's a lot better if you split those
- 13:47things. On to the next concept, which is
- 13:50load balancing. Now, going back to our
- 13:52client-server model, we had a server and
- 13:54we end up having clients. But, the
- 13:55problem is we might get a lot of users.
- 13:58And so, you have all those clients
- 13:59making requests to a single server. Now,
- 14:02a typical Node.js server is pretty
- 14:04powerful. It can usually satisfy up to
- 14:062,000 to 10,000 concurrent requests with
- 14:10a well-optimized Node.js server. But,
- 14:12when you go beyond that, you might need
- 14:13different ways to scale it because a
- 14:15single server has its own limits. And
- 14:18the easiest way to scale it is by load
- 14:20balancing. And that means you basically
- 14:23will create identical instances of your
- 14:26servers and then have this component
- 14:27that's an application load balancer
- 14:29splitting traffic between them. Now,
- 14:31load balancing is a topic by itself.
- 14:33There's different ways to split traffic
- 14:35and different criterias. And as a
- 14:37front-end engineer, you don't need to go
- 14:38so deep into it. But, do make sure that
- 14:40you know about it. Ideally, you're even
- 14:42capable of setting up a small load
- 14:44balancer using traditional web servers
- 14:46like Nginx and Docker Compose, you can
- 14:48have this set up on your local machine.
- 14:51Anyhow, the important thing is that you
- 14:52can reason your way through it. Most
- 14:54cloud providers like AWS or Google
- 14:56Cloud, they allow you to provision a
- 14:58load balancer in seconds. So, don't
- 15:00worry about it. It's very uncommon that
- 15:02you need to manually set up one, but
- 15:04it's important that you know about it.
- 15:06Now, congrats on making it this far.
- 15:07Make sure you subscribe so you don't
- 15:08lose any updates in the future, and
- 15:11let's move on to our next concept, which
- 15:13is container systems. So, we basically
- 15:15distributed our architecture into micro
- 15:17front-ends and different microservices,
- 15:20but the problem is that deploying all
- 15:22this would be a headache. Now we need to
- 15:25build up and provision infrastructure
- 15:27and pipelines, and they all might use
- 15:29different technologies. You might have a
- 15:31Python microservice and then a Node.js
- 15:33one. You might have a Vue.js application
- 15:36or a Next.js application. And this is
- 15:37all very complicated to get to
- 15:39production. So, in order to standardize
- 15:41the deployment, we can actually use
- 15:44Docker. And Docker is technology that
- 15:46allows you to package your application
- 15:48to Docker image. And the way you do that
- 15:50is that you take your code, and then you
- 15:53have a Dockerfile where you kind of
- 15:54declare your recipe of how we should
- 15:56package that code. And basically, based
- 15:59on that, you'll create a Docker image.
- 16:00The Docker image will contain all the
- 16:01application code and then the runtime.
- 16:04Let's imagine you're using Next.js, then
- 16:06it will contain Next.js, will have the
- 16:08runtime, which is Node, and then the
- 16:09operating system, which is usually
- 16:12Linux. And that is a full Docker image,
- 16:14and the advantage there is that whenever
- 16:16you find a host that runs Docker, you
- 16:18can run that image. You don't need to
- 16:20worry about the Node.js version or if
- 16:22they need to install PHP and all the
- 16:24dependencies that your application has.
- 16:25It really comes packaged all together.
- 16:28This is Docker image. You run it, open
- 16:30this port, you have a front end running.
- 16:31You don't need to know about what's
- 16:33inside. And this is wonderful for DevOps
- 16:36teams because all of the sudden they can
- 16:37take all those images and push them into
- 16:40a container orchestration system. A
- 16:42container orchestration system usually
- 16:44has a container deployment pipeline,
- 16:46which will run these containers. And so,
- 16:48running a container is not as easy as it
- 16:50sounds. You might need load balancing.
- 16:52Uh you might want to run several
- 16:54instances in parallel and be able to,
- 16:57you know, if a container fails, spin up
- 16:59another one really fast. So, all this
- 17:01kind of hard work it's built into
- 17:03systems like Kubernetes. You probably
- 17:05saw it in job office. Now, do you need
- 17:06to know Kubernetes as a front end
- 17:08engineer? No. But you do need to
- 17:10explain, you do need to know about it,
- 17:13and you will see a lot of front end
- 17:14positions that mention either Kubernetes
- 17:17or container systems or a ECS, which is
- 17:20the AWS alternative to Kubernetes, in
- 17:22the job description. Don't be scared.
- 17:24You don't need to become a DevOps
- 17:25engineer by tomorrow, but you should be
- 17:27able to, at a high level, understand
- 17:29where it fits in your architecture.
- 17:31Finally, a more front-end concept, a
- 17:33CDN, a content delivery network. So,
- 17:36what is a CDN? So, going back to a
- 17:37client-server model, imagine you are a
- 17:39client, you want to go on a website, you
- 17:42usually go to the server and get some
- 17:44static files. Static JavaScript and CSS,
- 17:47and you download that and you run that
- 17:49on your web browser. Let's imagine in a
- 17:51more hypothetical case that you are a
- 17:53user from the US and you want to visit
- 17:55the application that is based in Europe.
- 17:57For you to get the JavaScript and CSS
- 17:59and all the HTML, you have to go all the
- 18:01way to the Atlantic Ocean and come back.
- 18:03And that young trip adds latency, and
- 18:05there's no physical way you can work
- 18:07around it. No matter how performant your
- 18:10application is, there is the limit of
- 18:12the speed of light, because data only
- 18:14travels as fast as the speed of light,
- 18:16and over long distances, light is very
- 18:18fast, but it still will add around,
- 18:20let's say, 200 ms to 250 ms on every
- 18:23request of latency. So, a solution to
- 18:26that is to your client decrease the
- 18:27distance between you and where those
- 18:29files are hosted. And the easiest way is
- 18:31to use a content delivery network. And
- 18:33so, basically, a CDN will be a network
- 18:35of these edge locations that are placed
- 18:37all over the world, and what happens is
- 18:39that the server will push the static
- 18:41assets there, and whenever you make a
- 18:43request, you'll be redirected to the
- 18:45edge location that is closest to you.
- 18:46Now, this mechanism on how exactly are
- 18:49you getting redirected to that edge
- 18:51location that is closest to you, it's
- 18:52very interesting, and I do want to make
- 18:54a video about it, but I'm not sure if
- 18:56it's something you're interested in. If
- 18:58you want a video about it, let me know
- 18:59in the comments. It's a bit more
- 19:01technical, and it's not something you'll
- 19:02get in interviews. But, to be honest,
- 19:04it's really interesting. So, if you want
- 19:06me to make a video about it, just give
- 19:07me an excuse by letting me know in the
- 19:08comments, and I'll go ahead and do that.
- 19:10Now, a CDN is a distributed cache. So,
- 19:13the technical term for whenever you get
- 19:15the asset from the CDN, it's a cache
- 19:16hit. And whenever the server pushes a
- 19:19new version, that's called cache
- 19:21invalidation. And this concept is very
- 19:23closely related to what we call cache
- 19:25busting, which is a mechanism that
- 19:27module bundles use to invalidate your
- 19:30assets. So, we make sure that when you
- 19:31deploy a new version, users really get
- 19:34the latest one, not a previous version
- 19:35that is probably still sitting somewhere
- 19:37in the CDN. I do have other videos in
- 19:40the channel talking about it, so I'm not
- 19:41going to go deep into it. But make sure
- 19:43you're able to relate those things
- 19:44across the stack. Now, keep in mind that
- 19:46a CDN is the fastest and most
- 19:48cost-effective to increase web
- 19:49performance by serving optimized assets
- 19:52with the right cache policy out of the
- 19:53box. Because nowadays, CDNs do a lot
- 19:56more than just placing the asset close
- 19:58to the client. They also compress it,
- 19:59and they also take care of the caching
- 20:01policy. So, it's a really cheap way for
- 20:04you to fix most performance problems.
- 20:07Now, let's talk about design systems.
- 20:08And going back to our micro front-end
- 20:10discussions, we have a feature team,
- 20:12that's a product feature team, and then
- 20:14we might have the payments feature team,
- 20:15and they work separately and release
- 20:18independently. Those are completely
- 20:20separate product teams. And the problem
- 20:22there is that you might get code
- 20:23repetition or what we call visual
- 20:26divergence. You can have this silos
- 20:28mentality. And so, basically, a button
- 20:30in the product page will look
- 20:32differently than a button on the payment
- 20:34page. And so, you will lose what we call
- 20:35visual coherence, and people will start
- 20:37to notice that those are actually
- 20:39separate front-ends because they look
- 20:40differently. And that's not good from a
- 20:43product perspective, but it's also not
- 20:45good from a technical perspective
- 20:46because you have too much repeated code.
- 20:48And so, the solution is to have a design
- 20:49system. And in a design system, the
- 20:51first thing you do is to define your
- 20:52design tokens. That's basically your
- 20:55team information. What's your primary
- 20:57color, what is your border definition,
- 20:59your fonts families, and so on and so
- 21:02forth. And the modern way to do that is
- 21:04you add them as CSS custom properties,
- 21:07basically variables, at the global level
- 21:09in your CSS, probably in the shell micro
- 21:11front end, and then it gets fed into all
- 21:13the other micro front ends. But, you can
- 21:15go a bit farther and build reusable
- 21:16components that then the feature teams
- 21:18can use to assemble their features. So,
- 21:21you could build inputs and buttons, and
- 21:23basically they would consume that UI
- 21:25package and use it in all these
- 21:26different micro front ends. So,
- 21:28basically what you achieve is that the
- 21:29UI looks consistent, but also you don't
- 21:32repeat your code. The cool thing about
- 21:33the design system is you can take care
- 21:35of accessibility, you can have all those
- 21:37components unit tested, and basically
- 21:39you are applying at an architectural
- 21:41level the do not repeat yourself dry
- 21:44principle. And going back to AI coding,
- 21:46a solid design system really makes the
- 21:48difference between generating some
- 21:50component slob and having inconsistent
- 21:52styles and bugs that need a lot of
- 21:54rework, and really having a consistent,
- 21:56reliable coding agent output. Trust me,
- 21:59I've been building a lot of AI in the
- 22:01last couple of months, and the first
- 22:02thing I do when I make a new project is
- 22:04to tell to extract from whatever the
- 22:06design is the design system, because
- 22:08then you feed that into your coding
- 22:11agent through different sessions, and
- 22:13you still get consistent output. If you
- 22:14don't do that, the agent will make
- 22:16things up, and your UI would just look
- 22:18different, and it'll be obvious that it
- 22:20was live coded. Now, our next concept is
- 22:22a design to code MCP, that's a model
- 22:24context protocol server. And so,
- 22:26basically, you remember we have our
- 22:28design system, and normally you would
- 22:29import that to start building with it.
- 22:31But, nowadays, let's imagine that your
- 22:33designers built your design system
- 22:34initially in Figma, and then you
- 22:37implement it in your library. You can
- 22:38use the Figma MCP server with a coding
- 22:41agent to very quickly assemble features
- 22:45into the feature teams, into the
- 22:46vertical slices of your product. And
- 22:48this is the kind of workflow that most
- 22:49companies are moving towards. So, if you
- 22:51are a front end engineer, you got to
- 22:53make sure that you know how to use an
- 22:55MCP server, you know what an MCP server
- 22:57is, and ideally, whatever design tool
- 23:00your team is using, you can plug an MCP
- 23:02into that or you might have to even
- 23:03build that connection yourself. And then
- 23:06plug that into your AI coding agent like
- 23:08Cloud Code, which is the most used in
- 23:10enterprise or Codex. Now, a quick note
- 23:13on CSS architecture and design tokens.
- 23:15What you can do to take things even
- 23:17further is to apply the Atomic CSS
- 23:19methodology and with your design tokens,
- 23:21create atomic classes that you can use
- 23:23in your code base. So, basically, you
- 23:25define what the border would be, but
- 23:26then you create a dot border class that
- 23:28has that property. And then the only
- 23:30thing that other developers have to do
- 23:32is to use that class. And this is how
- 23:34Tailwind CSS works. So, keep this in
- 23:37mind. Tailwind CSS is an implementation
- 23:40of the Atomic CSS architecture style for
- 23:43CSS. And there's three more
- 23:45architectural styles, which I won't get
- 23:46into right now, but who knows, maybe
- 23:49I'll make a video about it later on.
- 23:51Next front end system design concept,
- 23:53the monorepo. So, basically, our
- 23:55applications right now are sitting in
- 23:56all these different repos where you have
- 23:58a GitHub repo for the shell, one for the
- 24:00payment micro front end, inventory micro
- 24:01front end, and it's all spread out into
- 24:04thousands of repos. And the problem
- 24:06there is that you end up with different
- 24:07code styles, different dependencies, and
- 24:09different quality standards. Some people
- 24:10might be using TypeScript, they might be
- 24:12using different linter configurations,
- 24:14and it's so easy to have what we call
- 24:16architectural drift or code style drift
- 24:19where two projects diverge too much.
- 24:21What's the problem with that? If you're
- 24:22a developer and you change teams, you
- 24:24have to relearn everything. So, you
- 24:27don't really leverage standardization.
- 24:29And the way to is to put everything into
- 24:31a single repository. So, basically, you
- 24:33have a big repository that contains all
- 24:35the other small applications, and you
- 24:37have tooling that works with both of
- 24:39them. So, for example, when you run NPM
- 24:40run build in a monorepo, all your
- 24:43applications will build individually.
- 24:45Now, when it comes to AI coding, a
- 24:47monorepo gives coding agents the context
- 24:49they need to make changes across service
- 24:51boundaries. So, for example, if you're
- 24:53building a micro front and you realize
- 24:56that you run into a reusable use case.
- 24:58You can easily extract that and propose
- 25:01it as a component into your design
- 25:02system. That might be a different repo.
- 25:04And you can do that with the coding
- 25:05agent in a single session if you have
- 25:08everything in a mono repo. If not, you
- 25:10have to somehow feed those two different
- 25:11repositories to your coding agent and
- 25:13everything becomes harder. So, combining
- 25:16micro frontends and microservices with a
- 25:18mono repo, it's usually the most
- 25:20efficient way to work with AI. Now, one
- 25:22thing I'm really excited about is also
- 25:24MC PUI. So, MC PUI is basically
- 25:27combining the traditional website with
- 25:29an LLM-powered app, and that's basically
- 25:30the glue in between them, the duct tape,
- 25:33let's say that. And so, basically, in a
- 25:35chat application, you usually send a
- 25:37chat query, and then that goes to the
- 25:39LLM, and the LLM will answer back, but
- 25:41in this case, it can actually answer
- 25:42with UI components. It can actually
- 25:44render products, for example, in the
- 25:46answer, not only text. But to do that,
- 25:49it has to somehow talk to the backend,
- 25:51and then we need to render some
- 25:53components. And that is what MC PUI
- 25:55solves. And this is an example I found
- 25:57in a recent website where I was looking
- 25:58for venues for events because we are
- 26:00organizing our annual in-person meetups
- 26:02at the Senior Dev, so we're going to
- 26:03meet all the engineers we work with
- 26:04around Europe or in the US, and I was
- 26:06looking for different cool venues where
- 26:08we could actually meet up. And so, I was
- 26:10talking to this chat UI, and all of a
- 26:11sudden, after a couple of questions, it
- 26:13started rendering those venues, and it's
- 26:15asking me for input, and then it would
- 26:17go into a deep search and give me even
- 26:19more venues. And as you see there, it
- 26:21can even render a map. And all this
- 26:23happens in a chat application. And I
- 26:25think this is the direction front end
- 26:26will go with LLMs, where we will
- 26:28integrate LLMs in web applications would
- 26:30be. A lot of people were saying that
- 26:32there's no more need for UI now that we
- 26:35have a chat. I disagree. I think the UI
- 26:37is a very useful way to communicate
- 26:39things, and I think you need front end
- 26:41developers, but we'll be able to combine
- 26:43the LLM approach with the traditional
- 26:45web approach. So, this component was
- 26:47rendered because the front end parsed it
- 26:50from the answers of the LLM. At a high
- 26:53level, the way this works is that the
- 26:55model harness will provide in the
- 26:57context the tool registries and all the
- 26:59MCP servers and then as a user prompt,
- 27:01and the LLM will send an answer back to
- 27:03the tool UI with the text answer, but
- 27:06also with an instruction to render a
- 27:07certain div, and then the tool, the web
- 27:10application has to follow that and
- 27:12render it.
- 27:13And just to really bring this to the
- 27:14code, the way we declare a resource or
- 27:17an MCP UI is we basically tell the LLM,
- 27:19"Hey, there's this resource, and you can
- 27:21use it in this case scenario, and this
- 27:24is the answer you can give." And then
- 27:25our front end has to take that and parse
- 27:27that and then render it. If you want me
- 27:29to make an in-depth video on how exactly
- 27:31MCP UI works, let me know, but this is
- 27:34one of those patterns that if you know
- 27:36really well, you'll really stand out
- 27:37because I believe it goes way beyond the
- 27:39current AI hype and it's actually
- 27:41something very useful for specific edge
- 27:44cases, and you'll see a lot of
- 27:45applications actually implementing this
- 27:47hybrid solutions. Now, to wrap up this
- 27:50talk a bit about performance, and the
- 27:52most important pattern that you'll ever
- 27:54see out mentioned in job descriptions
- 27:56for front end engineers, it's the Core
- 27:58Web Vitals. And so basically the Core
- 27:59Web Vitals are three metrics that
- 28:02quantify the three dimensions in which
- 28:04we measure the performance of a website,
- 28:07and those are the loading speed, the
- 28:09interactivity speed, and the visual
- 28:11stability. And the three Core Web Vitals
- 28:13that measure this are the Largest
- 28:14Contentful Paint, the Interaction to
- 28:15Next Paint, and the Cumulative Layout
- 28:17Shift. Basically, those are measured by
- 28:19Google, and they tell us what a good
- 28:21number would be. So, when you look at
- 28:23the Largest Contentful Paint, it would
- 28:24be the time it takes from when you hit
- 28:27enter to when the largest element in the
- 28:29web page is rendered. The Interaction to
- 28:31Next Paint, it's slightly different, and
- 28:32it has to do to when you interact. So,
- 28:34you do something, and then when do we
- 28:36repaint the UI? And the CLS is basically
- 28:38how much the UI changes when it loads.
- 28:41And so, the LCP and CLS have to do with
- 28:43the initial render, and the INP has to
- 28:45do with the re-renderings, which is very
- 28:47important when you talk about component
- 28:49frameworks. Now, to understand those,
- 28:50you need to understand the critical
- 28:52rendering path, which is all the steps
- 28:54you go from downloading some HTML to
- 28:56actually showing something on the
- 28:58screen. And that involves building the
- 29:00DOM, and then building the CSSOM, then
- 29:02building a render tree, then computing
- 29:04the layout tree, which is basically a
- 29:06tree of where all your nodes are and the
- 29:09positions and the width, and then
- 29:11transforming that into what we call
- 29:13paint operations that goes through to
- 29:15the GPU, then going to the composite
- 29:17phase, which has its own complexity, and
- 29:19I'm not going to talk about it, but just
- 29:21to summarize this, all those steps will
- 29:23happen, and then you have some
- 29:25re-renders because we use component
- 29:26frameworks, you get data, you start
- 29:28re-rendering, you finally finish, and
- 29:29then that's when you paint the LCP. So,
- 29:31all that time is measured in the LCP,
- 29:33and the bottom line here is, if you're
- 29:35shipping a lot of JavaScript, if you're
- 29:37shipping a lot of CSS, if you have to
- 29:39fetch a lot of data, and your server is
- 29:41slow, your application will be slow, and
- 29:43you'll get a very poor score in the LCP.
- 29:45The other thing that will happen if your
- 29:46application is poorly optimized is that
- 29:48you'll have layout switching all the
- 29:50time as you start loading things because
- 29:52CSS comes too late, and then fonts come
- 29:54in, and then some data comes in, and the
- 29:56browser will take screenshots of that
- 29:57and try to figure out if you are moving
- 29:59things too much. That would be the
- 30:00cumulative layout shift. And finally,
- 30:02the interaction to next paint is
- 30:04basically whenever you have a user
- 30:05event, and you have to go to the
- 30:07reconciliation and re-render after state
- 30:09update in your framework, and then you
- 30:11start again, you modify the DOM, and
- 30:14that triggers a re-paint. So, all that
- 30:16time it's quantified as the interaction
- 30:19to next paint. So, basically, if you
- 30:21have very slow re-renders or you're
- 30:23re-rendering too many components when
- 30:25users do something, then you have a very
- 30:27slow IMP. You can measure those metrics
- 30:30like house, and they go a lot deeper
- 30:31into this video on this channel about
- 30:33it, so make sure you check out that one.
- 30:35Our next concept also has to do with
- 30:36performance, and it's called splitting.
- 30:38Now, traditionally, a module bundler
- 30:40will take all our JavaScript and put it
- 30:42together in a single big file. But
- 30:44loading that single file will totally
- 30:47mess our core web vitals because we load
- 30:50too much JavaScript. So code splitting
- 30:53allows us to split our JavaScript to
- 30:55where it's being needed so we can really
- 30:57ship only the JavaScript that's needed
- 30:59to a specific page. And the easiest way
- 31:01to code split is by out. So basically,
- 31:03you would only ship to the slash login
- 31:05page the components that are
- 31:06specifically needed for the login. And
- 31:08if you have a dashboard that's very
- 31:10heavy with a lot of graphs, for example,
- 31:12you don't ship all that. So you
- 31:14selectively ship your JavaScript to
- 31:15where it's being needed instead of
- 31:16putting it together all in a single
- 31:18file. And all this is achieved with a
- 31:20module bundler like Webpack or Vit,
- 31:23which understands your bundle, splits
- 31:25it, and then dynamically loads it based
- 31:28on the path you're on, working together
- 31:30with your application router. And the
- 31:31higher overarching mental model here is
- 31:33lazy loading, which is the opposite of
- 31:35eager loading. Eager loading means you
- 31:37are really loading everything when a
- 31:39user lands on the page. And lazy loading
- 31:42means you load things as they needed. So
- 31:44for example, you might load certain
- 31:45things on scroll, or you might load
- 31:48certain things when they visit a certain
- 31:49page, or you might load certain things
- 31:51when they click on something. So you
- 31:52kind of wait for that user interaction,
- 31:54and when they land on the page, you
- 31:56really only ship the things they need.
- 31:58And all this is done in order to make
- 32:00this core web vitals better and working
- 32:02across the critical learning path. And
- 32:04finally, let's talk about rendering
- 32:05strategies. And nowadays, we usually
- 32:07work with modern component frameworks
- 32:09like React, Vue, and Angular. And the
- 32:11problem with those is that when you land
- 32:12on the page, you see a white screen. And
- 32:14the reason for that is because, you
- 32:15know, you load that empty HTML, and
- 32:17until you don't really run whatever
- 32:19render function they have, you don't
- 32:21really see much on the screen. That's
- 32:23what they call client-side rendering. So
- 32:26in an SPA architecture, in a single-page
- 32:28application with client-side rendering,
- 32:30you'd go, get your static file, and then
- 32:32you need to go get some dynamic data,
- 32:34and then finally render. And all that
- 32:36takes a lot of time. So, if you want to
- 32:38have a very performant website or a
- 32:40website that's ready to be crawled by a
- 32:42search engine, then client-side
- 32:44rendering is not the best choice for
- 32:46you. An alternative to this, if you have
- 32:47a static website where there's not a lot
- 32:49of interactivity, is to pre-render it on
- 32:51the server and ship it already rendered.
- 32:53So, when your client goes to get the
- 32:55static files, they already get HTML CSS
- 32:57and they don't need to run so much
- 32:59JavaScript. Now, again, this only works
- 33:01with static sites. If your static site
- 33:04will change often, let's say you have a
- 33:06blog and you want to publish new
- 33:07articles, then what you can do is
- 33:09incremental static generation. That
- 33:11means you only regenerate the pages that
- 33:14have changed. So, basically, your CMS
- 33:16will trigger a rebuild when you add a
- 33:18new blog post and that will get the
- 33:20dynamic data and go to a build pipeline
- 33:22and regenerate the only portion of the
- 33:24static files that change. So, it's a bit
- 33:26of a partial rebuilding of the website.
- 33:29The advantage here is not only that
- 33:30makes the build faster, but if I'm a
- 33:32client and I already downloaded part of
- 33:34your CSS and JavaScript, but I don't
- 33:35need to download re-download only the
- 33:37parts that change. Now, in most cases,
- 33:40this is built into a framework like
- 33:42Next.js, so you never have to worry
- 33:44about this yourself. And finally, we
- 33:46have server-side rendering. And in
- 33:48server-side rendering, the client will
- 33:49make a request to the front-end server,
- 33:51but then the front-end server will
- 33:52request the back-end server and get some
- 33:55data and then render the application on
- 33:57the server and then send it back
- 34:00pre-rendered. So, the client receives a
- 34:03full HTML page. So, you don't have this
- 34:05problem of the white screen. The problem
- 34:07you still have is that this page is not
- 34:10interactive yet because you haven't went
- 34:12through building the virtual DOM and
- 34:14attaching, for example, in the case of
- 34:15React, your virtual DOM to the actual
- 34:17DOM pre-built. And that's why you need
- 34:18to hydrate. And so, basically, you
- 34:20render that HTML page and then you have
- 34:23to execute your JavaScript, create
- 34:25internally the virtual DOM and then that
- 34:27virtual DOM is attached to the existing
- 34:29HTML markdown. That's what we call
- 34:30hydration. And finally, when you're
- 34:32hydrated, you might have to do some
- 34:34extra data fetching, so you might still
- 34:36need to go to the back. And again, this
- 34:38is one of the most complex approaches,
- 34:40so be careful with it. It's only useful
- 34:43whenever you need very, very fast
- 34:45performance or you need SEO. And a lot
- 34:48of people and a lot of companies jumped
- 34:50into this and they're doing server-side
- 34:51rendering, but that's like building an
- 34:53F1 car to go to the groceries. It's
- 34:56over-engineering and it creates a lot of
- 34:58problems and then everything gets
- 34:59slower. There's so many issues that will
- 35:01appear when you have this kind of setup.
- 35:03Technology to make it work is very
- 35:04complex. Bugs are harder to solve, so
- 35:07you want to stay away from it. And I
- 35:08really prefer simple solutions unless
- 35:11your use case really needs the highest
- 35:13performance. Oh, finally, let's talk
- 35:15about data fetching and I know we are
- 35:16front-end engineers, but it's very
- 35:18important that you also can work at a
- 35:19data layer. As I said before, we are
- 35:21moving towards front-end engineers being
- 35:22more full stack. And that is server-sent
- 35:24events. And all this has to do with
- 35:27real-time communication. When it comes
- 35:28to real-time communication that is not
- 35:31following the request-response cycle
- 35:33that I showed until now, you basically
- 35:35have to give it to it. You can do
- 35:37polling, you can do web sockets or
- 35:38server-sent events. And so polling would
- 35:40mean that you keep calling a specific
- 35:42endpoint until something happens. So,
- 35:44let's say we have a transaction that's
- 35:46processing, I can keep calling the
- 35:48status endpoint until it becomes
- 35:51completed. That's very easy to do with
- 35:53plain JavaScript and set timeout.
- 35:54There's nothing complex about it. The
- 35:56problem is you're calling your server
- 35:57way too much. And so it doesn't scale
- 35:59really well. You can have race
- 36:01conditions. So, it's a very simple
- 36:03solution, but not the most performant
- 36:05one. The next alternative would be web
- 36:06sockets, where you open a channel
- 36:09between the server and the browser and
- 36:10you can send updates and they can send
- 36:12you updates back. The problem here is
- 36:14that again, it's a lot of overhead and
- 36:16it's extremely good whenever you have
- 36:18bi-directional communication. Like
- 36:19you're working for a chat application,
- 36:20for example, where both the client and
- 36:22the server would keep sending chunks of
- 36:24messages. Now, for most use cases, this
- 36:27is not really needed and again, it's
- 36:29complex and it's very intense on the
- 36:31server. It needs a lot of resources.
- 36:33Now, with AI, we do need real-time
- 36:36communication, but it's only one
- 36:37direction. Because usually when you send
- 36:39a query to a chat application, you then
- 36:41just wait and they start sending you
- 36:43tokens back. So, it's the server sending
- 36:45a lot of messages, but you usually only
- 36:47send one. So, there's a lot of asymmetry
- 36:48between the client and the server. And
- 36:50the way to make this happen is with
- 36:52server-sent events, where you send a
- 36:53text message and that will create a
- 36:55conversation and on that endpoint,
- 36:58you're able to receive updates from the
- 36:59server. So, you receive all these
- 37:01tokens, but you don't send so many. And
- 37:03this approach is the one used by most
- 37:05LLM applications. If you've ever used
- 37:07the OpenAI NPM package in a React
- 37:09application to build a chat app with an
- 37:11LLM, that's exactly what they use under
- 37:13the hood. And you can actually figure
- 37:14this out if you go to your network when
- 37:16you use ChatGPT or Claude and find that
- 37:19conversation request and you'll see that
- 37:20the answer to that it's all these event
- 37:23streams. So, you're getting chunks of
- 37:25the answer. And that's how you build
- 37:26these cool UIs for LLMs. It's not
- 37:29WebSockets and it's not pulling, it's
- 37:32the server-sent event API. Make sure you
- 37:34look it up because it is something that
- 37:36will help you build product software
- 37:38products, software applications with AI
- 37:40embedded. Thanks so much, folks. If you
- 37:42want to go even deeper, make sure you
- 37:43check out these two videos on our
- 37:45channel. One of them is about all the
- 37:47front-end architectures that you need to
- 37:48know as a front-end engineer and the
- 37:49other one is about in-depth concepts
- 37:52about web performance, CSS, component
- 37:55frameworks that you really need and that
- 37:57show up in interviews at the senior
- 37:59level. And I'll see you in the next one.
About this transcript
This page contains the full transcript of Frontend System Design Explained w/ Senior Engineer (Microfrontends, Monorepo, MCP UI, Reactjs) by theSeniorDev, generated from the public captions YouTube serves with the video. The transcript has 7,923 words across 1,160 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.