DevOps Live Interview | DevOps | Cloud | DevSecOps | SRE #interview #devops #aws #docker #cloud — Transcript
Full transcript
- 0:04Hi Jadeja.
- 0:05>> Yeah, Akash.
- 0:08>> So, let me open your resume once. Just
- 0:10give me a sec.
- 0:13Okay.
- 0:15So,
- 0:16till then till then please start by
- 0:18introducing yourself.
- 0:20>> Yeah. So, my name is Jadeja Rajpalsinh
- 0:22and I'm a DevOps engineer with 7 year
- 0:24plus experience in designing, building,
- 0:26operating scalable and highly available
- 0:28cloud platform.
- 0:29My core experience lies in implementing
- 0:31the key pillars of DevOps that is
- 0:33automation, continuous integration
- 0:34delivery, infrastructure as a code,
- 0:36monitoring, and security.
- 0:38Throughout my journey, I have worked on
- 0:40multiple cloud platform, but most
- 0:43extensively in the side of AWS and a bit
- 0:45of GCP and Azure. I have worked on
- 0:47designing cloud native platform,
- 0:49Kubernetes based environment, and
- 0:51end-to-end CCD ecosystem for enterprise
- 0:53application.
- 0:54And if I talk about on the side of
- 0:56governance, then I have implemented
- 0:58secure multi-cloud architecture,
- 1:01centralized access management solution
- 1:03aligned with compliance standard like
- 1:04GDPR and HIPAA. I also been involved
- 1:07into the client discussion and pre-sales
- 1:09activity where I need to translate the
- 1:11business requirement into the technical
- 1:13solutions.
- 1:14Along with the hands-on engineering, I
- 1:16have certification like AWS certified
- 1:18solution architect, Red Hat certified
- 1:20system administration, and Red Hat
- 1:22Ansible automation exam. So, I think
- 1:24yeah, this is just a brief about me.
- 1:27>> Okay.
- 1:28So, currently in your organization, are
- 1:30you working independently or as a part
- 1:32of DevOps team?
- 1:34>> Um like
- 1:36we are a service-based company, so right
- 1:38now I'm working on two projects
- 1:39simultaneously. In one project, I'm
- 1:41working as an individual contributor
- 1:43where I independently manage the DevOps
- 1:45activity. And in another project, I am
- 1:47leading a team of three DevOps engineer
- 1:49where I'm responsible for task planning
- 1:51and technical guidance.
- 1:53>> Okay. Okay.
- 1:55So, let's start with the questions.
- 1:58Suppose you need to host a new
- 2:00application that is expected to handle
- 2:03around 5 million concurrent users.
- 2:06What How would you design that? And what
- 2:09all services and tools would you use?
- 2:12>> Okay, so 5 million concurrent users. So,
- 2:14before designing the architecture, I
- 2:15would first understand the application
- 2:17requirement, whether it will be
- 2:19stateless, stateful. We need to also
- 2:22understand the expected traffic pattern,
- 2:24the database requirement, compliance,
- 2:25and latency expectation. And based on
- 2:28that, I would design a highly available
- 2:29and secure platform. So, like if I talk
- 2:31about the upper layer architecture, then
- 2:33it would be a it would have a route 53
- 2:36for DNS DNS, then we we would have
- 2:39CloudFront for CDN and caching. We could
- 2:41use a web and shield for security.
- 2:44Then to distribute the traffic, we can
- 2:45use load balancer. And if you're going
- 2:47with the containerization approach, then
- 2:49either we can go ahead with ECS or EKS.
- 2:51I would choose ECS if it like if the
- 2:55architecture is not that much complex,
- 2:57and ECS has lower operational overhead.
- 2:59But if you want some sort of advanced
- 3:01feature of Kubernetes like
- 3:03GitHubs, some sort of custom networking,
- 3:06or advanced deployment strategy, then in
- 3:08that case, I would use EKS one. And also
- 3:10at the same time, we can use Argo CD for
- 3:13the deployment.
- 3:14And for the backend, we can go ahead
- 3:16with Aurora and DynamoDB for the
- 3:19database, Redis for the caching, S3 for
- 3:21the static content, SQS Kafka for
- 3:24asynchronous processing. And there are
- 3:26multiple tools for the monitoring and
- 3:27CCD that we can set it up over there.
- 3:29So, it usually depends on the budget.
- 3:31Like if the client can't like can't pay
- 3:33it, then they can't have it. So, we need
- 3:34to understand the cost factor as well
- 3:36there over there. So, but we can go
- 3:38ahead with some sort of open source
- 3:40tools like Prometheus, Grafana,
- 3:42CloudWatch. We can go ahead with some
- 3:43sort of CCD tools like Jenkins, GitHub
- 3:46Action, GitLab CI. Yeah, so I think this
- 3:48would be some of my checklist that I
- 3:50would follow.
- 3:52>> Okay. Cool. Suppose you successfully
- 3:56build this application and your
- 3:58application is already in the running
- 3:59state, right? And you what you notice is
- 4:02sudden traffic spikes up to 10x, right?
- 4:06How would you determine whether it's the
- 4:08genuine traffic or a DDoS attack?
- 4:11>> Okay. So,
- 4:13as a debugging point of view, firstly I
- 4:15would analyze the access log to look at
- 4:17the source IP, request pattern,
- 4:19geolocation, and the request path. If
- 4:21the traffic is coming from diverse real
- 4:23users with normal behavior and valid
- 4:24application request, it's likely a
- 4:26genuine traffic spike. So, over there I
- 4:28would simply scale up the infrastructure
- 4:30using the auto scaling. But, if I see a
- 4:32large number of repeated requests from a
- 4:34suspicious IP or maybe some unusual user
- 4:37agent, and also at the same time if they
- 4:39are requesting at or targeting a single
- 4:41endpoint, so I would suspect a DDoS
- 4:43attack over there. So, to mitigate it, I
- 4:45would use a AWS WAF, AWS Shield, or like
- 4:48we can use some sort of rate limiting
- 4:49functions over there, and it will block
- 4:51the malicious IP. Like, also at the same
- 4:53time, we need to continuously monitor
- 4:55the application to ensure the legitimate
- 4:58users are not getting impacted over
- 4:59there.
- 5:01>> Makes sense, yeah.
- 5:03Um okay. So, suppose as you mentioned
- 5:07auto scaling in your answer, that
- 5:09reminds me of Kubernetes. So, have you
- 5:12worked on Kubernetes, yes?
- 5:13>> Yeah, yeah. Actually, we worked on the
- 5:15Kubernetes side.
- 5:16>> Okay. That's cool. So, suppose you have
- 5:19three name spaces in a Kubernetes in the
- 5:21same Kubernetes cluster,
- 5:23how would you prevent the dev name space
- 5:25from communicating with the prod name
- 5:27space?
- 5:28>> Okay. So, like in our case, ideally we
- 5:31create a different cluster for like
- 5:33different cluster for different
- 5:35environment.
- 5:36But, let's say if we are having all the
- 5:38environment into a single cluster, then
- 5:40over there we could use the Kubernetes
- 5:42network policy to restrict the cross
- 5:44name space communication. By default,
- 5:46like I would deny all the ingress and
- 5:48egress traffic, then explicitly we can
- 5:50allow only the required communication
- 5:52within the namespace that is required
- 5:54over there.
- 5:56>> Got it. Got it.
- 5:57So, have you created any, you know,
- 6:00deployments or services or ingress out
- 6:02there from the YAML file? Have you ever
- 6:05>> like we are creating it from the scratch
- 6:06even though we are working into the helm
- 6:08chart format, so yes.
- 6:10>> Okay. So, just let me know what happens
- 6:13internally when you run a kubectl apply
- 6:16hyphen f the supposed deployment.yaml.
- 6:18What happens internally?
- 6:20>> So, whenever we apply the kubectl
- 6:21command, so basically kubectl reads the
- 6:23YAML for manifest, then it sends the
- 6:25request to the kube API server, and then
- 6:28API server validates the request, some
- 6:30sort of authentication and
- 6:31authorization, and then it store the
- 6:33data into the etcd.
- 6:36Then the scheduler assign the port to a
- 6:38suitable worker node, then the kubelet
- 6:40on that specific node receives the
- 6:41instructions and ask the container
- 6:43runtime to create the container.
- 6:45Finally, the port start running and the
- 6:46container, like, sorry, the controller
- 6:48continuously ensure that the desired
- 6:50state is maintained, like, whatever the
- 6:51replica that you are mentioning inside
- 6:52the deployment file.
- 6:54So, usually this is how the flow works.
- 6:57>> Okay. So, based on your previous answer,
- 7:01um I assume that you're pretty much well
- 7:05versed with the Kubernetes architecture,
- 7:07no? Like, you have Have you heard about
- 7:10etcd on the master plane?
- 7:12>> Yes, yes. So, basically it's a database
- 7:13where it store the data of everything.
- 7:16>> Okay. So, like,
- 7:19answer this, like, why did Kubernetes
- 7:22choose etcd instead of any other
- 7:24database? They could have chosen any
- 7:25database, but why etcd?
- 7:29>> Um
- 7:30like, to be honest, like, I haven't
- 7:32looked into that depth, but my
- 7:34understanding is that etcd is built for
- 7:35distributed system, and it provides a
- 7:38consistency and leader election,
- 7:39specially the leader election.
- 7:41So, which are important for Kubernetes,
- 7:42so that is what my current
- 7:44understanding, but I would
- 7:45uh need to check into a bit more. Yeah.
- 7:49>> So, yeah, fine, fine, fine. Makes sense.
- 7:51So,
- 7:52>> [clears throat]
- 7:53>> your your you must have hosted, you
- 7:55know, microservice uh applications or
- 7:57the websites out there, right?
- 7:59>> Mhm.
- 8:00>> So, what do you think like how does the
- 8:03request reach your pod when a user
- 8:04accesses your website?
- 8:07>> So, like whenever user sends a request
- 8:08to our website, the request first
- 8:10reaches our reaches to the DNS. So, it
- 8:13can be route 53, it can be GoDaddy, it
- 8:15can be anything, and then it resolves
- 8:17the domain from there. And then it
- 8:18reaches like if you are using the
- 8:20application load balancer, then it will
- 8:21firstly go to the application load
- 8:23balancer in our case.
- 8:24Then it forwards the request to the
- 8:25Kubernetes ingress controller. And the
- 8:27ingress controller the ingress routes
- 8:29the request to the appropriate
- 8:31Kubernetes service, and then the service
- 8:33forward it it to the one of the healthy
- 8:35pod using the kube proxy.
- 8:38>> Okay. Okay.
- 8:40Uh let me just allow me a sec. Let me
- 8:42have a look at your resume.
- 8:45Okay. So, in your achievement section,
- 8:49you mentioned that you migrated 500 TBs
- 8:52of S3 bucket of data of S3 bucket from
- 8:55one AWS account to another.
- 8:58Mhm. Um like from the point of
- 9:01accomplishments, isn't that
- 9:02straightforward? You you could have just
- 9:05used AWS data sync, right?
- 9:08Can you just elaborate on that?
- 9:10>> So, even though to perform that
- 9:11migration activity, we used the AWS data
- 9:13sync. So, the migration part was never a
- 9:16bigger challenge for us. The real
- 9:18challenge was to how to very re-verify
- 9:21that all the data has been migrated
- 9:22successfully without missing any object
- 9:24because since the data was into the 500
- 9:26TB of data, so it was very difficult for
- 9:28us to re-verify the objects whether
- 9:30everything has been successfully
- 9:31migrated or not. So, to solve that
- 9:34specific issue, I I the some more
- 9:36services of AWS like I used the S3
- 9:38inventory. It basically generates a CSV
- 9:40format containing all the objects in a
- 9:42in the bucket. And I then loaded that
- 9:44inventory into a DynamoDB table and like
- 9:47then simply compared the source and the
- 9:49destination inventory to identify any
- 9:51missing object with the help of that
- 9:52specific DynamoDB database.
- 9:55>> Okay. Okay. Makes sense.
- 9:58Um okay. So, how like uh in your 6-7
- 10:02years of DevOps career, I hope that you
- 10:06must have built a lot of pipelines out
- 10:08there, right?
- 10:09>> Correct.
- 10:10>> Uh so, can you just walk me through the
- 10:12best CI/CD pipeline that you built?
- 10:16>> So, actually we I worked on multiple
- 10:17pipelines with different types of tools
- 10:19and different different type of
- 10:20technology, but in my recent project
- 10:22that currently where I'm working on, so
- 10:24let me walk you through over there. So,
- 10:26basically the approach would be almost
- 10:28similar for almost all the pipeline all
- 10:29the tools. So, usually whenever a
- 10:31developer pushes the code, the pipeline
- 10:33will automatically get triggered. It
- 10:35will firstly pull the latest code from
- 10:36the Git. It will run some unit test
- 10:38cases. Then it will perform some sort of
- 10:40code quality analysis using the
- 10:41SonarQube. Then while creating the
- 10:43image, we also use some sort of
- 10:45vulnerability scanning tool like Trivy
- 10:46and everything. Then it pushes the image
- 10:48to the ECR.
- 10:50Then for the CD part, we are having
- 10:52different repositories. So, over there
- 10:53we update the image tag in the GitHub
- 10:55repository.
- 10:57Then Argo CD automatically detect the
- 10:59change and it synchronize it with the
- 11:01our Kubernetes cluster.
- 11:03And then the application was uh in our
- 11:05in our current project, the application
- 11:06was deployed using the rolling update
- 11:08strategy. And post deployment, health
- 11:10checks were performed before making
- 11:11before marking the deployment as
- 11:13successful.
- 11:15>> Uh you mentioned uh the application was
- 11:17deployed using rolling update, right?
- 11:20>> Correct.
- 11:21>> Uh so, why didn't you use a canary
- 11:23deployment and instead you used rolling
- 11:25update? Because canary deployment is too
- 11:27much in fashion, no?
- 11:28>> Correct.
- 11:29>> So, why?
- 11:30Yeah.
- 11:30>> So, basically the thing is that we are
- 11:32at the stage where we are enhancing the
- 11:34infrastructure.
- 11:35So, like in the current project, like we
- 11:37are using the rolling update because
- 11:39like the application was backward
- 11:40compatibility, and the deployment risk
- 11:42was very low. So, it provided us a
- 11:44zero-downtime deployment with minimal
- 11:45operational overhead.
- 11:47But let's say if we have to set up a
- 11:48canary deployment for some sort of
- 11:51business-critical application or major
- 11:52releases, we could really
- 11:54implement the canary deployment. It just
- 11:56like we need to just
- 11:58do some sort of basic testing. Like the
- 12:00basic implementation, it will just allow
- 12:01us to allow give us a feature to
- 12:03release a new version to a small
- 12:05percentage of user. Then we can simply
- 12:07monitor the key matrices like the error
- 12:08rate, latency, and then we can simply
- 12:10gradual gradually increase the traffic.
- 12:12And if any issue is detected over there,
- 12:14then we can simply roll back into the
- 12:16previous version. So, if if needed, then
- 12:18we can also implement the canary
- 12:19deployment. But we are at the stage
- 12:21where we are enhancing the
- 12:22infrastructure.
- 12:24>> Got it. Got it. Makes sense. Yeah.
- 12:27So, like while creating CI/CD
- 12:30deployments out there, there you might
- 12:32have faced a lot of challenges.
- 12:34Sometimes some havoc, some
- 12:37issue in the configuration. So, let me
- 12:39give you a scenario. Let's say a
- 12:41developer accidentally commits an AWS
- 12:44access key
- 12:45and the secret key to the Git Git
- 12:47repository.
- 12:49What And that that's a blunder, no? What
- 12:52immediate actions would you take and how
- 12:55would you prevent this from happening in
- 12:57future?
- 12:58>> So, firstly, what I will do, I will just
- 13:00simply disable or rotate the
- 13:02compromised access key so that it cannot
- 13:04be used anymore. And then I will go to
- 13:07the AWS, and then we can we could check
- 13:09the cloud trail to just to identify if
- 13:11any suspicious activity has been done
- 13:13with the those compromised key or not.
- 13:16And then
- 13:17for the future, we can integrate some
- 13:19sort of Git secret or Git leak into the
- 13:21CI pipeline to detect whether any of the
- 13:23secret is getting pushed or not before
- 13:25the code is getting merged.
- 13:26And then like at the like the best
- 13:28practices is we should store some sort
- 13:30of sensitive credential into the AWS
- 13:33secret manager or we could save it to
- 13:34HashiCorp Vault as well.
- 13:36And also, at the same time, we need to
- 13:38educate the developers on secure
- 13:40credential management as a part of the
- 13:42development workflow.
- 13:44>> Okay. Okay, makes sense.
- 13:47So, sup- let's suppose
- 13:51like by using your CI/CD pipeline or
- 13:54something, uh you deployed a new version
- 13:56of the application, right? The
- 13:58deployment is showing successful, that
- 14:00means the green deployment, right? So,
- 14:02it's green on your side, but the
- 14:04customers are, you know, getting a lot
- 14:06of 502 bad gateway errors.
- 14:09Right? So, how would you troubleshoot
- 14:10that?
- 14:12>> So, firstly,
- 14:14uh okay, the deployment goes green, but
- 14:15still the users are impacting or getting
- 14:17502 errors.
- 14:18>> Yeah, 502. Yeah, yeah.
- 14:19>> So, maybe firstly, I will check the logs
- 14:21of the pod, and maybe I will check
- 14:23whether that pods are up and running or
- 14:25not. Maybe I would also check some sort
- 14:26of readiness and liveness probe as well.
- 14:29And then, I will verify whether the
- 14:30service is correctly pointing or
- 14:31forwarding the traffic to the healthy
- 14:33pods or not.
- 14:34Then, even though if that is also
- 14:35correct, then I will check the ingress
- 14:37controller or the application at the LB
- 14:39logs to identify whether the from where
- 14:41the 502 is being generated. Then, also I
- 14:44will at the same time, I will I will try
- 14:45to check some sort of application logs
- 14:47for any sort of startup or runtime
- 14:49error.
- 14:50Then, like, I will also ensure that the
- 14:52the application is listening on the
- 14:53correct port and the service target port
- 14:56matches the container port that we
- 14:58defined inside the deployment file.
- 15:00And even though if the issue is still
- 15:02persistent over there, so my first
- 15:04preference would be to roll back it to
- 15:05the
- 15:06previous stable version to restore the
- 15:08services while investigating the root
- 15:10cause. And we can we could do it onto
- 15:12the lower environment. So, firstly, I
- 15:13will roll roll back it to the previous
- 15:15version, so that at least the end
- 15:16customer would not impact anymore.
- 15:19>> Got it. So, like, suppose as you
- 15:22answered, you will check this you will
- 15:24check a lot of things. So, like during
- 15:27your investigation of the checks, you
- 15:30find that the suppose you during your
- 15:32investigation of the checks, you find
- 15:33that the 502 error is caused by the
- 15:36database, right? And the database CPU is
- 15:39consistently above 90% since you made
- 15:44the deployment. And one there's one
- 15:46query that is being executed hundreds of
- 15:48times per minute because of the because
- 15:52of the particular application
- 15:53functionality that's assume. So, how
- 15:55would you solve this?
- 15:57>> So, the firstly I would do is to confirm
- 16:00whether the repeated query is actually
- 16:01necessary or it's an application issue.
- 16:04Let's say if the data doesn't change
- 16:05frequently like whenever user is
- 16:07requesting that specific query. So, I
- 16:09would introduce Redis as a caching
- 16:11layer. So, instead of hitting the
- 16:12database directly for every request, the
- 16:14application firstly it would go to the
- 16:16Redis and if the data is available it
- 16:17would return the response directly from
- 16:20there. If the data is not present in
- 16:21Redis, it would query the database and
- 16:23then store the result in Redis for the
- 16:25subsequent request. So, like this would
- 16:27significantly reduce the number of
- 16:30database queries. It will also reduce
- 16:32the CPU utilization. So, maybe this is
- 16:35how we could improve the response time
- 16:37for the request.
- 16:39>> Okay.
- 16:41Okay.
- 16:42So, have you worked on Terraform?
- 16:44>> Yes, I'm working on Terraform. Yeah.
- 16:46>> Okay. So, how would you structure
- 16:49Terraform code for large enterprise
- 16:50project with multiple environments like
- 16:53dev, prod, and stage?
- 16:56>> So, we can create a modular structure in
- 16:58which like we can create separate module
- 17:00folder for component like VPC, EKS, RDS,
- 17:03or whatever the services that we are
- 17:05using into our current project. And then
- 17:07there would be a separate environment
- 17:08folder such as dev, stage, prod in which
- 17:10where each environment will call the
- 17:12same module with the different variable
- 17:14values.
- 17:15So, it would like also at the same time
- 17:16each environment would have its own back
- 17:18end and state file to keep the interest
- 17:20isolated. And then we could simply
- 17:22integrated that Terraform with a CCD
- 17:24pipeline to so that every infrastructure
- 17:26changes goes through review, validation,
- 17:28and approval being before being getting
- 17:30applied over there.
- 17:31So this is what we can do here.
- 17:33>> Okay. So while working with Terraform,
- 17:36Terraform being too sensitive and you
- 17:39know, you you need to do your work in a
- 17:41good fashion, right? So
- 17:45you must have faced some kind of you
- 17:47know, issue while issue as in like
- 17:50suppose I'll give you an example. Like
- 17:52what like what would happen if someone
- 17:54manually changes the infrastructure in
- 17:56AWS outside of Terraform? What would
- 17:59what like what happens in that case?
- 18:01>> Let's say
- 18:02if they made some sort of changes into
- 18:03the AWS cloud and if we are not having
- 18:05that code into our Terraform, then
- 18:07whenever we will try to run the
- 18:08Terraform plan, the Terraform will
- 18:10detect the drift and shows the
- 18:11difference between the desired state and
- 18:13the current state. So like it depend on
- 18:15the change. Like if the manual change is
- 18:16not part of the Terraform code,
- 18:17Terraform apply may revert it back to
- 18:19the desired state. But let's say if the
- 18:21manual change is needed into our current
- 18:23code right now. So I would firstly
- 18:25review it and then update the Terraform
- 18:26code so that the infrastructure and the
- 18:28code remains in sync. So this is what
- 18:30this is how we can solve it.
- 18:33>> Okay.
- 18:34So
- 18:35when do you use depends on? Can you just
- 18:37give me an example?
- 18:39>> So like
- 18:40I use depends on when I want to make
- 18:43one resource is created only after
- 18:46another resource is ready. So there is a
- 18:47dependency. Over there we could get So
- 18:49if I give you a small example, then
- 18:51let's say we create a private route
- 18:53while creating VPC, right? So I want to
- 18:56ensure that the net gateway because
- 18:58inside the route we attach the net
- 18:59gateway to the routes. So I want to
- 19:01ensure that the net gateway is fully
- 19:02created before the route points to it.
- 19:04So in such case I can accidentally use
- 19:06the depends on to guarantee that the
- 19:08creation order.
- 19:10>> Okay. Okay.
- 19:12So let's come back to AWS maybe just a
- 19:15few questions on AWS.
- 19:17What would happen if an entire AWS
- 19:21region goes down and how would you
- 19:23recover your application?
- 19:25>> Like we are already deploying our
- 19:27application with the help of
- 19:28multi-region. But even though in our
- 19:31current project since we are working
- 19:32with US based client, so we perform some
- 19:35sort of disaster recovery drills every 3
- 19:37months to ensure our recovery process
- 19:39actually works. So, there are multiple
- 19:41steps that we need to follow. So, we are
- 19:42keeping a backup of our database like we
- 19:45use the database snapshot. It will
- 19:47automatically copy like it will
- 19:49automatically create the snapshot and
- 19:50there is lambda function which will move
- 19:51the snapshot from one region to another
- 19:53one another one, sorry. And then we also
- 19:56enable the S3 cross region application
- 19:57that is simply CRR so that our
- 19:59application data is already available in
- 20:01the DR region. And since our
- 20:03infrastructure is managed through
- 20:04Terraform, so we can quickly recreate
- 20:06the infrastructure in the secondary
- 20:07region. That won't be an issue much
- 20:08more.
- 20:09And once the application is up and
- 20:11ready, like we can simply update the
- 20:12route 53 to route the traffic to the
- 20:14healthy region.
- 20:15And we have already tested this DR
- 20:17process multiple times in our current
- 20:19project. So, like right now like we are
- 20:22much confident that we could cover the
- 20:24application like we can recover the
- 20:25application within the agreed RTO or the
- 20:28RPO instead of figuring it out at the
- 20:30run time.
- 20:31>> Okay.
- 20:33Okay, okay.
- 20:34I was just going through your resume. I
- 20:37find one thing very interesting.
- 20:40You worked on
- 20:43>> [clears throat]
- 20:43>> HIPAA and GDPR compliances, right?
- 20:46>> Correct.
- 20:47>> So, what was your exact role in this?
- 20:50Can you just mention?
- 20:51>> So, basically I was working for a client
- 20:53and they are into the health care
- 20:54sector. So, since they are applying
- 20:56around the data of the patient. So, they
- 21:00need to be a bit more
- 21:02concerned about the security part. So,
- 21:04like currently they They into the Europe
- 21:06under the GDPR compliances. So, when
- 21:07they expanded their business to the US
- 21:09so they also needed to comply with the
- 21:11HIPAA requirement. So, as a part of the
- 21:13initiative, our our focus was on
- 21:15strengthening the infrastructure from a
- 21:17security and compliance perspective. So,
- 21:20if I give you a small example, there
- 21:21were a lot a lot of checklist over
- 21:23there, but if I give you a small
- 21:24example, then uh we ensured the private
- 21:26S3 bucket were accessed through VPC
- 21:28endpoint instead of the public internet.
- 21:30Then we are implementing some sort of
- 21:31governance policy in our community
- 21:32cluster to uh to enforce the secure
- 21:35deployment uh standards. Then we make
- 21:37sure that the PI and the PHI data was
- 21:39never written to the application logs by
- 21:42masking or removing the sensitive
- 21:43information from that logs.
- 21:45Then we are also like we also enforce
- 21:47the encryption. We also use some sort of
- 21:49least privileges I am access and
- 21:51regularly we review the security and
- 21:53everything, yeah.
- 21:55>> Okay. Okay.
- 21:57So, let [clears throat] me give you a
- 21:59real-life scenario, right?
- 22:01>> Mhm.
- 22:01>> So, uh right now FIFA is uh too much
- 22:05trending and the ads and the banners in
- 22:07the FIFA. You must have heard that the
- 22:09in different countries they show
- 22:11different kind of ads or different kind
- 22:12of pages, right? So,
- 22:15>> [clears throat]
- 22:15>> like I would like to question on that in
- 22:17CloudFront, if you want to show
- 22:19different pages based on the user's
- 22:21country, how would you implement that?
- 22:25>> Um different pages based on the their
- 22:28country, right?
- 22:29>> Yeah.
- 22:30In CloudFront.
- 22:31>> Oh, CloudFront, right.
- 22:32So, there would be multiple ways of
- 22:34doing it. Uh like the firstly, the
- 22:36simplest would be like we can write up a
- 22:37small logic inside the application that
- 22:40will handle the uh
- 22:42logic which will uh move the traffic
- 22:44from one page to different one according
- 22:46to the user's country. But let's if you
- 22:48are specifically going with the
- 22:50CloudFront, so there is a there is one
- 22:51more option that we can do it uh
- 22:53CloudFront
- 22:55provides us the CloudFront function.
- 22:57So, we could use the CloudFront
- 22:58functions over there and there is a uh
- 23:00header like we can use uh, viewer
- 23:02country header, like I don't exactly
- 23:04remember the header, but there is some
- 23:05sort of CloudFront viewer country, some
- 23:07sort of header like that.
- 23:07>> Yeah.
- 23:08>> And then we can write up a small
- 23:09CloudFront function, and that it will
- 23:10automatically inspect that specific
- 23:12header, and rewrite the request URI or
- 23:14the route the users to a different uh,
- 23:17content over there. Let's say if the
- 23:18user is coming from India, so it will
- 23:20uh,
- 23:21access on the server page like with some
- 23:23sort of landing page. So we could create
- 23:25a CloudFront function. Then there are
- 23:27multiple ways of doing it. So even
- 23:28though we could implement some sort of
- 23:29Lambda function through which
- 23:30Lambda@Edge which will handle this
- 23:32specific logic. So there are multiple
- 23:34ways of doing it.
- 23:36>> Mhm. Okay.
- 23:38Okay.
- 23:40So
- 23:41uh, have you worked in multi-cloud
- 23:42environments?
- 23:44>> Yeah, so like uh, like within this uh, 7
- 23:47or 6 years of my cloud journey, so I got
- 23:50a chance to work on different different
- 23:51projects. I worked on AWS, GCP, Azure. I
- 23:54did a migration of like multiple times I
- 23:56did a migration of from GCP to AWS as
- 24:00well. I did a lot of uh, I have created
- 24:03a lot of pipeline in terms of Azure
- 24:04DevOps as well. So like it's it's like
- 24:06they have worked on almost all the cloud
- 24:08provider, but most extensively on the
- 24:10side of AWS.
- 24:12>> Most ex- And I also see that you have
- 24:16uh,
- 24:17AWS certification, right?
- 24:18>> Correct. Correct. So yeah, it's a
- 24:19professional one.
- 24:21>> Yeah, it's a professional one. So when
- 24:23did you get this?
- 24:25>> So I did that in the year somewhere
- 24:27around 2023. So maybe it would be in
- 24:30some sort of expiration state or maybe
- 24:32like it would be a bit closer to the
- 24:34expiration.
- 24:35>> Yeah, it's like fine. It's like
- 24:38Okay.
- 24:39Uh, let me go through your resume once
- 24:42more.
- 24:43Just allow me a second.
- 24:45Okay. So I was Yeah, I was asking that
- 24:49Mhm. Like, [clears throat] can you just
- 24:51tell me more about Kibana, like a little
- 24:53about Kibana?
- 24:55>> So right now we have used the Kibana in
- 24:56our current project, so it is helping us
- 24:58to uh,
- 25:00strengthening the security of the cabin.
- 25:01So, let me give you a small example. So,
- 25:03we have like it's just a YML file in
- 25:04which you could create any sort of
- 25:06policy that you wanted to use in your
- 25:07Kubernetes cluster. In our case, like
- 25:09I'm using like this is one of the policy
- 25:11that we are using.
- 25:12So, that we have mentioned the account
- 25:14ID of the AWS account account AWS
- 25:15account ID. So, that whatever the image
- 25:18that is present into that specific AWS
- 25:20account ID, only those deployment would
- 25:22be able to deploy. They won't they won't
- 25:24be able to deploy any sort of basic like
- 25:26Ubuntu, Nginx, Python. They won't be
- 25:28able to deploy anything. Only the images
- 25:30that is present on that specific AWS
- 25:32account ID. So, this is how we are
- 25:34restricting to use
- 25:36secure images that is present on our AWS
- 25:38account. Yeah.
- 25:40>> Okay.
- 25:42Okay, makes sense. Um
- 25:45yes, I guess I covered most of the
- 25:47points.
- 25:49Um yes, so yep, I'm good from my side.
- 25:53Um maybe if you want to have some
- 25:54questions, please spam them.
- 25:56>> So, actually I don't have any specific
- 25:58question, but just wanted to understand
- 25:59that what kind of project or business
- 26:01domain that your organization is working
- 26:02on and also maybe some bit idea about
- 26:04the tools and the technology.
- 26:07>> Okay. So, in our organization, we work
- 26:11with multiple clients across multiple
- 26:12domains such as, you know, health care,
- 26:15fintech, a lot of e-commerce clients out
- 26:17there, retail clients out there. So, the
- 26:20technology stack varies like depending
- 26:23on the client based on the client
- 26:24requirement. But most of our project
- 26:26revolve around AWS, Kubernetes,
- 26:28Terraform, Docker, GCP, and monitoring
- 26:30tools.
- 26:31So, as a DevOps engineer, you'll likely
- 26:34work on different projects and, you
- 26:37know, get exposure to multiple
- 26:38technologies over the same time. Yeah.
- 26:41I guess I answered your question.
- 26:43>> Yeah, yeah. Yep, perfect. I think so,
- 26:45yeah, that is the only question that I
- 26:47have right now.
- 26:48>> Okay. Okay.
- 26:50So, nice meeting you, Jadeja. I'll just
- 26:53pass on my feedback to the HR. She'll
- 26:55get back to you.
- 26:55>> Sure. Thank you so much. Thank you so
- 26:57much. Yeah, have a nice day.
- 26:58>> Thank you.
About this transcript
This page contains the full transcript of DevOps Live Interview | DevOps | Cloud | DevSecOps | SRE #interview #devops #aws #docker #cloud by jadeja rajpal sinh, generated from the public captions YouTube serves with the video. The transcript has 4,957 words across 817 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.