DevOps Live Interview - Fintech - 35 LPA | DevOps | Cloud | #interview #devops #aws #docker #cloud — Transcript
Full transcript
- 0:01Hello, Jadeja.
- 0:03>> Yeah, hey Ravi.
- 0:04>> Hey, how are you?
- 0:06>> I'm really good. How are you?
- 0:08>> I am also good. So, let me open your
- 0:11resume.
- 0:13Uh meanwhile, you can introduce
- 0:14yourself. Okay?
- 0:16>> Yeah. So, basically, my name is
- 0:18Jadeja Rajpasingh and I'm a DevOps
- 0:19engineer with around 7 year plus
- 0:20experience in designing, building, and
- 0:22managing all sort of highly available
- 0:24cloud platform.
- 0:25So, my core expertise lies in
- 0:27implementing the key pillars of DevOps,
- 0:28that is automation, continuous
- 0:30integration delivery, infrastructure as
- 0:32a code, monitoring, observability, and
- 0:34of course, the security part. So,
- 0:35throughout my career, I have worked on
- 0:37multiple cloud platform, but uh
- 0:39extensively on the side of AWS and a bit
- 0:41of GCP and Azure.
- 0:43So, I'm good at designing the cloud
- 0:44native platform. I worked on
- 0:46Kubernetes-based environment and set up
- 0:48the end-to-end uh CI/CD ecosystem for
- 0:50the enterprise application.
- 0:51I have successfully contributed to
- 0:53several large-scale transformation
- 0:55project like including the migration of
- 0:57Kubernetes cluster from GKE to Amazon
- 0:58EKS, GKE to EKS. I have modernized
- 1:01monolithic application from micro-based
- 1:04into the microservice-based
- 1:05architectures.
- 1:06Then, uh on the governance side, I have
- 1:08implemented the secure multi-account
- 1:10cloud architecture, centralized identity
- 1:12and access management, and solution
- 1:14aligned with the complex standards such
- 1:15as GDPR and HIPAA. And uh apart from it,
- 1:18I've also been involved into the client
- 1:19discussion and pre-sales uh engagement
- 1:21where I need to uh get in touch with my
- 1:23client and need to convert their
- 1:25business requirement into the technical
- 1:26architectures. And uh throughout this
- 1:29journey, I did some sort of
- 1:29certification like uh AWS Certified
- 1:31Solutions Architect Professional, Red
- 1:33Hat Certified System Administration, and
- 1:35Red Hat Certified Ansible Automation.
- 1:37And beyond all these technical
- 1:38responsibilities, I have trained more
- 1:40than uh 50 plus engineers in my uh
- 1:42current organization.
- 1:44>> Okay, yeah.
- 1:46So, uh you mentioned in your
- 1:48introduction that you have trained over
- 1:5050 DevOps engineer.
- 1:52So, could you tell me more about your
- 1:54role in that?
- 1:56>> So, like basically, like I've worked at
- 1:59a service-based company. So, every year
- 2:01around 20 or 30 DevOps engineer join the
- 2:03company.
- 2:04So, as a part of their onboarding, I
- 2:05conduct the technical sessions covering
- 2:07core DevOps concepts like the basic of
- 2:09AWS, Kubernetes, CICD, Terraform, and
- 2:11many more tools. So, my responsibility
- 2:13is to help them help the new engineer
- 2:15build a strong foundation so that they
- 2:16can contribute effectively to the client
- 2:19project. So, that is what usually we
- 2:21follow a procedure into our current
- 2:22organization.
- 2:23>> Okay, yeah, yeah. Okay.
- 2:25So, in your resume you mentioned that
- 2:27you worked in a lot of migration
- 2:29project, right?
- 2:31>> Mhm, yes, correct.
- 2:32>> So, yeah, let me give you a scenario.
- 2:34So, your company has decided to migrate
- 2:37the application from EC2 to
- 2:40Kubernetes.
- 2:41So, how will you decide Kubernetes is
- 2:43the right choice?
- 2:46>> Mhm, so like before even though going to
- 2:49the Kubernetes, firstly I will try to
- 2:51understand the current application and
- 2:52the problem that you are trying to
- 2:53solve.
- 2:54So, we need to check the application
- 2:55compatibility whether the application
- 2:57can be containerized easily or not
- 2:59without having the major changes. And
- 3:02like after that like what I will do,
- 3:03I'll compare the benefit of the
- 3:04Kubernetes like the auto scaling,
- 3:07self-healing, and the zero downtime
- 3:08deployment. But along with that, we need
- 3:09to also
- 3:10like keep one factor in mind that is the
- 3:12operational complexity and the cost
- 3:14because if like whether the improved
- 3:16scalability and the automation like it
- 3:18should justify the additional cost as
- 3:19well.
- 3:20Then what I will do, I'll just simply
- 3:23do a small POC with a non-production
- 3:25application. And if the results are
- 3:26positive, then I will create a proper
- 3:28document and recommend like with moving
- 3:30to the Kubernetes.
- 3:33>> Yeah, okay, yeah. Sounds good. So, now
- 3:36you decided Kubernetes is a right choice
- 3:38after doing POC and you know, creating a
- 3:40document. But like how would you plan
- 3:43the migration to ensure there is a no
- 3:45downtime for end user?
- 3:48>> So, firstly, the first step would be
- 3:50that we need to containerize the
- 3:51application and test it thoroughly. Then
- 3:54I will set up the Kubernetes
- 3:55infrastructure. It can be on anything,
- 3:56EKS,
- 3:58Azure AK AKS or GK or anything. Along
- 4:01with that, we need to properly set up
- 4:03the networking, ingress, storage and
- 4:05everything. Then what I will do I will
- 4:06just deploy the application in a
- 4:07non-production environment and perform
- 4:09some sort of functional and security
- 4:11testing over there. Then I will set up
- 4:12the CI/CD pipeline for automated
- 4:14deployments. Then at the end I will use
- 4:17a gradual migration strategy like the
- 4:18blue/green or the canary deployment
- 4:20instead of switching all the traffic at
- 4:21once. Then I will monitor the
- 4:23application health, the logs, some sort
- 4:24of matrices during the migration.
- 4:26And at the end like I will keep a
- 4:28rollback plan ready so that if any
- 4:29issues occur at that time while doing
- 4:31the migration, traffic can immediately
- 4:33be rolled back to the EC2 instance. So
- 4:35this is what usually I will
- 4:36follow a procedure to migrate the
- 4:39application.
- 4:41>> Okay, yeah. So let's assume after
- 4:43migration, deployment becomes slower
- 4:46than EC2.
- 4:48So how would you troubleshoot
- 4:50that?
- 4:52>> Okay, so after the migration, right? So
- 4:53firstly I will try to
- 4:55identify where the delay is happening
- 4:57because Kubernetes itself doesn't
- 4:59necessarily make the deployment slower.
- 5:00So it can be multiple reasons. So
- 5:01firstly I would check at the level of
- 5:03CI/CD because inside the CI/CD there we
- 5:06use multiple stages like the image
- 5:07build, image push, deployment. So
- 5:09firstly I will try to analyze if
- 5:10something is going wrong or maybe if
- 5:12something is taking bit more time into
- 5:14the form of CI/CD. That is my first
- 5:16thing that is that is my first thing
- 5:18that I will look into it. Then I will
- 5:19verify if the images are sizes are too
- 5:21large, then I will try to optimize the
- 5:22Docker file using the multi-stage build
- 5:24and smaller base images. So that will
- 5:26ultimately help to reduce the deployment
- 5:28time. Then the other factor that can
- 5:31also cause the slow deployment into the
- 5:33form of Kubernetes that can be like the
- 5:35the pod scheduling. Some sort of
- 5:37sometimes the pod remains in the pending
- 5:39pending state because of insufficient
- 5:41CPU or memory or the node availability.
- 5:43Then inside the Kubernetes I will try to
- 5:45review the readiness and liveliness
- 5:47probe as well because incorrect probe
- 5:49configuration can also increase the
- 5:51delay time.
- 5:52Then at the end, I will try to monitor
- 5:54the cluster resource like I can enable
- 5:55some sort of cluster auto scaler so that
- 5:57if any nodes are going to the bottleneck
- 5:58or if any nodes is creating an issue, it
- 6:01will automatically get scaled up. So,
- 6:03usually this would be some of my
- 6:04checklist here.
- 6:06>> Okay, yeah. Okay. So, now during a
- 6:09deployment,
- 6:11all old pod were terminated before new
- 6:13pod become ready.
- 6:15So, that is causing down time. So, how
- 6:18would you prevent this?
- 6:20So, it's not happened in future.
- 6:22>> So, even though in in many of our
- 6:24project like we have faced the same
- 6:25issue. So, usually this happens due to a
- 6:27incorrect deployment strategy or you can
- 6:29say that some sort of incorrect
- 6:31configuration inside the Kubernetes. So,
- 6:33if you want to if you wanted to prevent
- 6:34such type of situation, then the first
- 6:36thing would be like we need to
- 6:38properly set up the readiness probes
- 6:39inside the Kubernetes so that only the
- 6:41traffic will reach to the healthy pods.
- 6:43And also like I will
- 6:45when I will use the rolling update
- 6:46strategy, so there we need to also
- 6:47mention the max available and the max
- 6:49surge value so that the pod will not get
- 6:51terminated until the new pods are
- 6:53getting ready over there.
- 6:54And there is one more policy that we
- 6:56could set up inside the Kubernetes
- 6:57cluster that is the pod disruption
- 6:59budget so that it maintains a minimum
- 7:01number of available pods during the
- 7:02deployment. And lastly like for the
- 7:05critical application deployment, we can
- 7:07go ahead with the blue green or the
- 7:09canary deployment like I usually I
- 7:10prefer the canary deployment to achieve
- 7:12the zero down time deployment.
- 7:14>> Yeah, it will avoid the down time. Yeah,
- 7:16okay. So, now you know, there is a
- 7:18situation where some application should
- 7:21not run together on the same node due to
- 7:24XYZ reason. So,
- 7:27so because uh
- 7:29sometime they consume too much CPU
- 7:32or
- 7:33So, how would you ensure proper
- 7:35placement across node?
- 7:39>> So, like like like, say if there are
- 7:41five pods for a same application or
- 7:43different type of application and we
- 7:44want each application to go into the
- 7:46different nodes, right?
- 7:47So, also, we can also put some sort of
- 7:49condition that no two pods get together
- 7:51into the pods, so we can use the pod
- 7:52anti-affinity over there. So, like, by
- 7:54doing this, like, our
- 7:56like, it will automatically,
- 7:58like,
- 7:59it will it will like, only if if one pod
- 8:01is going into one of the node, then pod
- 8:03anti-affinity will give us the option or
- 8:05the feasibility that no other pod would
- 8:06get scheduled into that specific node
- 8:08over over there.
- 8:09And if the application needs to run only
- 8:11one only on a specific node, then over
- 8:12there I'll label the nodes and then we
- 8:14can use the node selector or the node
- 8:16affinity over there into our
- 8:17configuration file. So, this is how we
- 8:19can do that.
- 8:21>> Okay. So, after migration, the this is a
- 8:24scenario like developer is requesting
- 8:27shell access to production pod for
- 8:29debugging.
- 8:30So, would you allow it?
- 8:32>> Oh, no, no, actually not at all because
- 8:34I would not allow any sort of direct uh,
- 8:36shell access to the production pod by
- 8:37default because it will introduce the
- 8:39security and the compliances.
- 8:40So, like, what I would encourage, like,
- 8:43I would encourage the developer to use
- 8:44the logs, to use the metrics or maybe
- 8:47whatever the tool that we are using,
- 8:48Prometheus, Grafana, DataDog, whatever
- 8:49the tools that we are using. So, maybe
- 8:51they can check the logs from there.
- 8:53But if still something is there, if the
- 8:55developer like, insisting me to give the
- 8:57permission if they really need it, so I
- 8:59will
- 9:00provide the temporary control access
- 9:01using the RBA with proper approval like
- 9:03from my management and everyone. Then
- 9:05also, along with the access, I will try
- 9:07to log and audit all the actions what
- 9:08they're performing with that specific
- 9:10access. [laughter]
- 9:11And after like, once the debugging is
- 9:13completed, then I will immediately
- 9:14revoke those temporary access over
- 9:16there.
- 9:17>> Okay, okay. Yeah.
- 9:18So, now now imagine your Docker registry
- 9:21become unavailable during a one of your
- 9:24production deployment.
- 9:26So, what impact would you expect after
- 9:28this incident?
- 9:31>> Okay, so image registry is going down.
- 9:32So, firstly, I'll check whether the
- 9:33required Docker image is already present
- 9:35on the worker node. And if the image is
- 9:37already there, the existing code will
- 9:39still run without any sort of impact.
- 9:41But if the Kubernetes is trying to
- 9:43deploy a new version, so the deployment
- 9:44will fail and the new code will not get
- 9:46started over there. So to minimize the
- 9:48risk, what I can do, I can simply use
- 9:51some sort of highly available or the
- 9:52private registry such as ECR so that we
- 9:54can also like at the same time we can
- 9:56also enable some sort of image caching
- 9:57where possible and we can also perform
- 9:59some sort of disaster recovery plan for
- 10:01the image
- 10:04uh that image ratio too so that if the
- 10:06primary region goes down,
- 10:08there would be a backup
- 10:09registry in another region so that we
- 10:11could we can pull the images from there.
- 10:13So that is what we can do easily.
- 10:15>> Okay, yeah.
- 10:16So now assume your one of your teammate
- 10:19accidentally deleted
- 10:21the production name space.
- 10:23Uh
- 10:24what will you do? Like Kubernetes name
- 10:26space. So in that case, what will you
- 10:28do?
- 10:30>> Okay, so they deleted the production
- 10:31name space only, right? So basically
- 10:32firstly I will stop the ongoing
- 10:34deployments or the changes to prevent
- 10:36the further issues.
- 10:37Then I will check if there is any sort
- 10:38of backup tool like generally we use the
- 10:40Valero in our Kubernetes cluster so that
- 10:42it takes the backup of everything like
- 10:44it takes the backups of name space,
- 10:45resource and everything. But let's say
- 10:47we don't have a backup. So if there's no
- 10:49backup exists, then I will simply
- 10:50redeploy the name space and workload
- 10:52using the GitHub repository because
- 10:53since we are deploying everything with
- 10:54the help of Argo CD. So we are having
- 10:56all the manifest file, we are having all
- 10:58the answer of each and every
- 10:59microservice. And then I will verify
- 11:01that all the applications are healthy
- 11:02and working like as expected after the
- 11:04restore.
- 11:05And then finally, I will try to
- 11:07investigate the root cause and implement
- 11:08some sort of
- 11:09safeguard like Rback, some sort of
- 11:11approval workflow and some sort of
- 11:13backup policy to prevent the similar
- 11:14incident so that no one would be able to
- 11:16delete the production name space again.
- 11:18>> Okay, okay. So let's say let's assume
- 11:21after migration is done, after 60 year
- 11:24after 6 month, management ask whether
- 11:27Kubernetes has delivered value. So in In
- 11:30case, what KPI would you present to
- 11:32management?
- 11:34>> Okay, so we wanted to compare whether
- 11:35the migration that we did that is
- 11:37successful or not, right?
- 11:38>> Yeah, right.
- 11:39>> So,
- 11:40there are multiple factors over there.
- 11:41So, the first thing that would be the
- 11:43deployment time. Like I will try to
- 11:44compare the before and the after the
- 11:46migration of the deployment time, like
- 11:47which one is taking the lesser
- 11:48deployment time. Then I will also try to
- 11:50check some sort of application
- 11:51availability, some sort of resource
- 11:53utilization, like when we are going with
- 11:54the Kubernetes, so whether it is
- 11:56optimizing the resources or not. Then
- 11:58the major like one of the major factor
- 11:59that is the infrastructure cost. Like I
- 12:02will try to compare the cost before and
- 12:04after the migration. Then also we need
- 12:06to check like how like sometimes the
- 12:07scale like sometimes there might be a
- 12:09chance that the costing has been
- 12:10increased because we are implementing
- 12:11some sort of application scalability
- 12:12feature as well. So, we need to also
- 12:14understand like how quickly that it
- 12:16handles the traffic spike. So, it
- 12:17depends on the business requirement as
- 12:18well. Then at the end I will try to like
- 12:21some I will try to check some sort of
- 12:23operational matrices like the deployment
- 12:24success rate, some sort of
- 12:26rollback frequency, some sort of MTTR,
- 12:29which is the mean time recovery. So,
- 12:31that are some of my KPIs that I would
- 12:33check before
- 12:34giving the conclusion to them.
- 12:38>> Okay.
- 12:39So, in your CI/CD, your appli- it is
- 12:42showing your application is deployed
- 12:44successfully. But for some user
- 12:47still see the old version.
- 12:50So, in that case, what could be the
- 12:52reason why some of some users are still
- 12:55seeing the old version of application?
- 12:58>> let's say if if you're working into the
- 13:00production environment and if the
- 13:01infrastructure is quite big. So, firstly
- 13:02I will check whether the deployment is
- 13:04like whenever we make the deployment, so
- 13:06sometimes it takes the time to roll out
- 13:08all the deployments all the ports. So, I
- 13:10will check whether all the ports are
- 13:11still in the progress or if all the
- 13:13ports are has been updated. And let's
- 13:15let's say if everything is updated, then
- 13:16I will verify the service like whether
- 13:17it is routing the traffic correctly to
- 13:19the new port or not.
- 13:20Then at the same time we need to check
- 13:22some sort of things into the side of
- 13:25browser, some sort of CDN, some sort of
- 13:26caching. If you're any sort of caching,
- 13:28then maybe we need to invalidate the
- 13:30caches so that it
- 13:32serve the new content.
- 13:34And one more thing, like if we are using
- 13:36the canary or the blue/green deployment,
- 13:38so it's expected that the some users see
- 13:40the old version and while others see the
- 13:42new version until that specific
- 13:44deployment goes complete or like whether
- 13:45that specific deployment get completed.
- 13:48So sometimes they like we can
- 13:50face some sort of issues over there.
- 13:53>> Mhm, okay. That's yeah, correct. So now
- 13:56how would you design your CI/CD pipeline
- 13:58so bad deployment never reaches to
- 14:00production directly?
- 14:03>> Okay. So like firstly I would say that
- 14:05we cannot guarantee that a faulty
- 14:07deployment will never reach the
- 14:08production.
- 14:09But yeah, we can reduce the chances by
- 14:11adding some sort of proper validation
- 14:12step inside the pipeline. So inside the
- 14:14staging like we can integrate some sort
- 14:16of unit test cases, some sort of
- 14:17security test cases inside the pipeline.
- 14:19Then before pushing the images we could
- 14:21scan the images so that if any
- 14:23vulnerability is there inside that
- 14:24image, we could get to know before
- 14:26making a deployment.
- 14:27Uh then
- 14:28like also I will not deploy any any sort
- 14:30of things or any deployment directly to
- 14:32the production. Firstly I will try to
- 14:33validate into the lower environment that
- 14:35is development or the staging
- 14:36environment. And like if required we can
- 14:39also integrate some sort of manual
- 14:41approval inside the production
- 14:42deployment as well.
- 14:43Then at the same time as already
- 14:45mentioned that we can go ahead with some
- 14:46sort of canary and the blue/green
- 14:47deployment so that it will minimize the
- 14:49risk. It won't commit like we can't give
- 14:51the commitment that the faulty
- 14:53deployment will never reach but at least
- 14:54we could minimize the risk by doing all
- 14:56such status of things.
- 14:58>> Yes, you have to add multiple gates on
- 15:00CI, right?
- 15:01>> Yep, correct.
- 15:02>> So now you know there you have multiple
- 15:04micro service in your application. Like
- 15:06it could be 50 or
- 15:0850 plus, let's assume.
- 15:10So for that would you create a separate
- 15:13CI/CD pipeline for each one?
- 15:16>> Um like
- 15:18to be honest, it depends on the project.
- 15:19But in most cases, yes, like I create a
- 15:22separate pipeline for each microservice.
- 15:23So, this allows a team to develop, test,
- 15:26and deploy their independent services
- 15:27without affecting the other one. But,
- 15:29let's say I will not duplicate the
- 15:30pipeline logic. I will create a reusable
- 15:32pipeline template, and so that each
- 15:34service can use the same template with a
- 15:35different configuration. So, with the
- 15:38help of this particular reusable
- 15:40pipeline template, it will keep the
- 15:41pipeline consistent and easier to
- 15:42maintain. And it follows a dry
- 15:44principle, which is don't repeat
- 15:45yourself. So, I think so, yeah, this is
- 15:47usually what I do.
- 15:49>> Mhm, separate pipeline because also, you
- 15:52know, rollback become easier, right?
- 15:55>> Correct. But, at the same time, if if we
- 15:57are using any sort of reusable
- 15:58component,
- 15:59then I will create a specific template
- 16:01for it, so that they can use those
- 16:03reusable components from that specific
- 16:04template.
- 16:06>> Okay, okay, yeah.
- 16:08So, now, like, you know, your company
- 16:10decided to migrate pipelines from
- 16:13Jenkins to GitHub Action due to XYZ
- 16:16reason, due to license expire, or
- 16:18whatever. So, how do you ensure a smooth
- 16:21migration without affecting existing
- 16:23deployments?
- 16:25>> So, firstly, we need to understand the
- 16:28existing Jenkins pipeline, the
- 16:29integration, the deployment process, the
- 16:31plugin, the users, however they're set
- 16:34up. So, firstly, I will try to
- 16:35understand the existing platform.
- 16:36Then, I will try to set up the same
- 16:38thing into the GitHub Action, then I
- 16:39will migrate one non-critical
- 16:40application
- 16:41just for a POC purpose. And then, I will
- 16:43run the Jenkins and the GitHub Action in
- 16:44parallel for both of like for some time,
- 16:46so that we could compare the result. And
- 16:48once the new pipeline is stable, then I
- 16:49will gradually migrate the remaining
- 16:51applications over there. And throughout
- 16:53the migration, I will closely monitor
- 16:55our deployment like and keep the Jenkins
- 16:56as a rollback option until the migration
- 16:58is fully completed. So, usually, this is
- 17:00what I will do.
- 17:01>> Okay.
- 17:03So,
- 17:04there is a scenario, okay? So, your
- 17:06application must be deployed
- 17:08simultaneously to
- 17:10multi-cloud, like, let's assume AWS and
- 17:13Azure.
- 17:14So, Uh, one deployment succeeded, but
- 17:17the other fails. So, how would your uh
- 17:20pipeline handle this scenario?
- 17:23>> Um, so firstly I will just stop the
- 17:25pipeline and identify where the why the
- 17:27Azure deployment getting failed. Then I
- 17:29will not mark the deployment as
- 17:30successful until both AWS and the Azure
- 17:33deployment gets
- 17:34completed successfully. And if they need
- 17:36if let's say if AWS is already running
- 17:38the new version, so firstly I will like
- 17:40I will either roll back it to the
- 17:42previous one or
- 17:44uh, like we depend like on the business
- 17:45requirement. So, if we can roll back it
- 17:47to the previous version, we can simply
- 17:48do it or else I will just keep it keep
- 17:50it as as it is running. And then I will
- 17:52try to fix the Azure issue and then
- 17:54rerun the pipeline and ensure that both
- 17:55environment are on the same new
- 17:57application version. And then at the end
- 17:59like I will try to monitor some sort of
- 18:00both deployment to verify whether they
- 18:02are healthy before completing the
- 18:03release. So, like it it will be a part
- 18:06of CICD pipeline. Like firstly the AWS
- 18:08pipeline will get triggered. Sorry, AWS
- 18:09pipeline will get deployed successfully
- 18:11then the Azure will get deployed. And if
- 18:13if both the deployments are getting
- 18:15successful, then I will make the
- 18:16deployment as successful otherwise I
- 18:17will just roll back for both of them.
- 18:19So, this is what we can do.
- 18:22>> Okay. So, do you have experience with
- 18:24Terraform as well or IAC tool?
- 18:26>> yeah. So, even though like for like from
- 18:28last 3-4 years I'm extensively working
- 18:30into the side of Terraform only whatever
- 18:31the infrastructure that I'm creating it
- 18:32up.
- 18:33>> Okay.
- 18:33>> Yeah.
- 18:34>> So, now your organization organization
- 18:36decided to manage all their
- 18:38infrastructure using Terraform only. But
- 18:41most of the resources were created
- 18:43manually. So, how would you approach the
- 18:46migration? Like yeah.
- 18:48>> Understood. So, the the resources are
- 18:50already there into the side of AWS,
- 18:51right? So,
- 18:52>> AWS. Yeah.
- 18:53>> Yeah. So, firstly I will try to identify
- 18:54and prioritize the existing AWS
- 18:56resources. I will try to I will write
- 18:57the Terraform code for those resources.
- 18:59And instead of recreating them, I will
- 19:01use the Terraform import.
- 19:04With help of import Terraform import you
- 19:05you can bring the existing resource into
- 19:07your Terraform state file. Then I will
- 19:09try to run the Terraform plan to verify
- 19:10that the Terraform doesn't propose any
- 19:12sort of unexpected changes. And once
- 19:14everything is validated, I will simply
- 19:15uh use the Terraform as a single source
- 19:17of truth for all the future
- 19:18infrastructure changes. So, we need to
- 19:20go into the repetition. Like like we
- 19:22will make the changes. Then again, we
- 19:23will run the Terraform plan. We need to
- 19:24go into the loop until and unless it
- 19:26doesn't show us any sort of
- 19:28uh
- 19:28uh changes over there. So, we need to
- 19:30repeat those changes again and again.
- 19:32>> Yeah, you have to run the Terraform
- 19:34import. Then you have to add the code on
- 19:36Terraform file.
- 19:38And then
- 19:38>> Correct.
- 19:39>> Repeat.
- 19:39>> to run the Terraform plan.
- 19:41And ultimately, I think it should share
- 19:43it should uh share share that like there
- 19:45is no infrastructure changes. Then only
- 19:47we would be able to 100% sure that it's
- 19:49working fine.
- 19:51>> Okay. Okay. So, now uh how do you
- 19:53prevent multiple engineers from making
- 19:55Terraform changes to the same
- 19:57environment simultaneously?
- 19:59>> So, it's a default feature of the
- 20:00Terraform. So, we can simply use the
- 20:02remote backend such as a S3 bucket to
- 20:03store the Terraform state file. And I
- 20:05would simply enable the state locking
- 20:06using the using the DynamoDB so that
- 20:08only one Terraform operation can run at
- 20:10a specific time. And also, I will ensure
- 20:13that all the Terraform changes goes
- 20:14through the CI/CD pipeline. Instead of
- 20:16being running it from the manually. So,
- 20:18inside the CI/CD pipeline, if someone
- 20:19has triggered something, then the second
- 20:21pipeline would automatically go into the
- 20:23queue. And also, at the same time, I'll
- 20:24separate the workspace or the state file
- 20:25by a for different different environment
- 20:27like the dev, QA, production. And also,
- 20:29like if required, we could also
- 20:30mention the code review and approval
- 20:32before applying the infrastructure
- 20:34changes over there.
- 20:36>> Okay. Yeah, so in multi-cloud scenario,
- 20:39so now uh like uh you need to provision
- 20:41infrastructure across multiple
- 20:43multi-cloud like AWS, Azure, GCP using
- 20:46Terraform only, okay?
- 20:48So, how would you organize the code base
- 20:50to maximize reusability or
- 20:52maintainability?
- 20:55>> So, firstly, I will just separate the
- 20:57code by like I will separate the code by
- 20:59cloud provider like the AWS, Azure, GCP
- 21:01while keeping a common repository
- 21:03structure over there. I will try to
- 21:04create a reusable Terraform modules over
- 21:06there so that for common services like
- 21:08the
- 21:09uh networking, some sort of security,
- 21:10some sort of computer storage, we could
- 21:12use the those usable Terraform modules.
- 21:15And then I will keep the environment
- 21:16specific configuration like the dev, QA,
- 21:17prod for using the different type of
- 21:19variable files.
- 21:21Then I will maintain some sort of
- 21:22separate state file for each cloud and
- 21:24each environment to avoid the conflict.
- 21:25Then like we already discussed like I
- 21:27will use the remote backend with the
- 21:29state locking and manage all the
- 21:30deployment with the help of CICD
- 21:31pipeline. So this is what I could do.
- 21:35>> Okay, yeah. With state lock locking and
- 21:37everything. Yeah, okay.
- 21:38So now uh
- 21:40how would you prevent a developer from
- 21:42accidentally deleting production
- 21:44resources using Terraform? Because you
- 21:46know, a Terraform code will now store on
- 21:49GitHub or version control system. So how
- 21:51would you prevent it? Yeah.
- 21:53>> So firstly, like we need to follow the
- 21:55least privilege principle. So I will
- 21:57restrict the production access using the
- 21:58IM role.
- 21:59Then along with that, I will also ensure
- 22:01that all Terraform changes goes through
- 22:03a CICD pipeline and mandatory code
- 22:05review and approval.
- 22:06Then inside the Terraform, we could use
- 22:08some sort of block like the prevent
- 22:10destroy is equal to two true. So that it
- 22:12will not let us delete the critical
- 22:14resource like the database, some sort of
- 22:16production cluster.
- 22:17Then I will also require developer to
- 22:18review the Terraform plan before they
- 22:21apply any sort of before they apply any
- 22:22sort of changes over there. And yeah, I
- 22:24think so this is And also we could
- 22:25separate out the production and
- 22:26non-production environment for the
- 22:28better practice. Yeah.
- 22:29>> You have You have to add life cycle
- 22:31policy.
- 22:32>> Life cycle policy. That would be the
- 22:33prevent destroy It's the life cycle
- 22:35policy over there. Yeah.
- 22:36>> Mhm.
- 22:37Get it. So now you know, have you worked
- 22:40with FinOps as well?
- 22:42Implemented FinOps in your
- 22:44organization before?
- 22:45>> Like whenever Yeah, usually we follow
- 22:47the like we worked into the side of
- 22:48FinOps as well where we need to automate
- 22:49the infrastructure cost and everything
- 22:51for our client.
- 22:53>> So now okay, let me give you a scenario.
- 22:54So now your company want to reduce cost
- 22:56by 40%. Let us assume AWS cloud, okay?
- 23:01>> Okay.
- 23:01>> But they don't want to they don't want
- 23:03any impact on application performance.
- 23:06So where would you start?
- 23:09>> So like I will analyze the current AWS
- 23:12cost using the cost explorer and
- 23:14identify the top cost consuming
- 23:16services.
- 23:17Then I would check whether the EC2
- 23:18instance EKS nodes RDS or storage are
- 23:21over provisioned and try to right size
- 23:23them. Okay.
- 23:25Then we need to also check like we we
- 23:26can improve the resource utilization by
- 23:28setting up some sort of proper condition
- 23:30like the CPU memory some sort of request
- 23:32limit into and we can also try to remove
- 23:34the ideal resources. If we are not using
- 23:36them it should not be there into our
- 23:37cloud. Then I will try to use some sort
- 23:39of auto scaling some sort of cluster
- 23:40auto scaling also if you wanted some
- 23:43sort of advanced
- 23:45scaling then we can use carpenter to
- 23:47ensure that only the resource that we
- 23:48actually need so with the help of
- 23:50carpenter provides us some sort of
- 23:51advanced feature over there. And for
- 23:53some sort of non-critical workloads like
- 23:55the dev QL Q Q environments so I we
- 23:58could also try to implement some sort of
- 24:00sport instance machine over there. And
- 24:02also we could go ahead with some sort of
- 24:03reserve plan and the saving plan as well
- 24:05because this is also one of the
- 24:07factor that we could consider and like I
- 24:09would continuously monitor some sort of
- 24:11cost and like utilize utilization to
- 24:13ensure that the optimization doesn't
- 24:15affect the application performance.
- 24:17>> Okay. Yeah.
- 24:20So now
- 24:21you know in one of your AWS account your
- 24:24primary AWS region become suddenly
- 24:26unavailable due to war or whatever.
- 24:30So how would your application recover in
- 24:33on that case?
- 24:34>> So even though right now I'm working for
- 24:36them like some of my clients and they
- 24:38are from US and Europe so over there we
- 24:39need to perform some sort of DR over
- 24:41there. So I will ensure that the
- 24:43disaster recovery setup is already there
- 24:44in place with a secondary AWS region. So
- 24:47if first region fails then the traffic
- 24:48will
- 24:49shift to the secondary region using the
- 24:50route 53 or there is another service
- 24:52that you can use that is the AWS global
- 24:54accelerator service. Then I will verify
- 24:56that the application database all all
- 24:58the other dependencies are healthy in
- 24:59the DR region. And if required, we could
- 25:01also restore the latest data from the
- 25:03backup or use the cross region
- 25:05replication for minimal data loss. And
- 25:07once the primary region is available
- 25:08again, then I will validate it and plan
- 25:11a controlled rollback over there.
- 25:12So, that is what we could do.
- 25:15>> Okay. So, in now in case of production
- 25:18alert, like severity one or P1 alert,
- 25:21so walk me through your first 30
- 25:24minutes. What action would you take on
- 25:26that case? Will you inform your manager
- 25:28or whatever action
- 25:30>> So, usually even though right now in our
- 25:32current
- 25:33>> project, usually we get an alert like
- 25:35there are some sort of SEV1, SEV2. That
- 25:37depends on the criticality criticality
- 25:39of that particular alert. So, we usually
- 25:41once I will get some sort of mail or
- 25:44some sort of like we usually have a
- 25:45portal where we would be able to see
- 25:47what are the sort of severity issues are
- 25:49happening in the production environment.
- 25:51So, firstly I will acknowledge the alert
- 25:52and assess the impact. Like is it I'll
- 25:55identify which application or the users
- 25:57are all the services are getting
- 25:58affected over there.
- 25:59Then I will simply join the incident
- 26:01bridge and inform the relevant team like
- 26:02the developer, infrastructure, the
- 26:04database team, whoever is involved over
- 26:05there. Then we'll check some sort of
- 26:07dashboard, some sort of logs, matrices,
- 26:09then try to check if any recent
- 26:11deployment to identify the possible root
- 26:13causes. And if a recent deployment
- 26:15causes the issue, then I will simply
- 26:17roll back it to the previous version to
- 26:19restore the services. And once the
- 26:20application is stable, I will continue
- 26:22investigating the root cause and monitor
- 26:24the system to ensure that the issue
- 26:26doesn't
- 26:27reoccur once again. And after the
- 26:29incident is resolved, I will conduct a
- 26:31root cause analysis that is RCA and
- 26:33implement the preventive measure on top
- 26:35of it.
- 26:37>> Okay. Yeah. Okay, so this is a last
- 26:39question from my side. Finally, I know
- 26:43finally if you decided to join our
- 26:45company as a senior DevOps engineer and
- 26:48you were given complete ownership of our
- 26:50platform,
- 26:52so what are the first three improvement
- 26:54you would look for before making any
- 26:57changes?
- 26:59>> Any changes, right? So, firstly, to be
- 27:01honest, like I will try to understand
- 27:03the current platform because I will also
- 27:05need some sort of time to review the
- 27:07architecture, CCD pipeline, the
- 27:09monitoring, security, infrastructure,
- 27:11infrastructure, and everything.
- 27:13Like there are many factors. Like you
- 27:14guys would be using monitoring,
- 27:16alerting, backup, disaster recovery
- 27:17plan. So, firstly, I will try to
- 27:19identify the current pain points over
- 27:21there and try to identify some sort of
- 27:23technical debt which is there in two of
- 27:24your existing platform.
- 27:27Then I will try to look for some sort of
- 27:29automation gap, some sort of cost
- 27:31optimization, some sort of performance
- 27:32bottleneck, and the developer experience
- 27:34improvement. I will try to check the I
- 27:36will try to take a feedback from the
- 27:37developer as well. Like are they happy
- 27:39with our system or not? Then I will
- 27:41prioritize changes based on the business
- 27:43impact instead of making the unnecessary
- 27:45modification. So, it usually depends
- 27:46like how the infrastructure has been set
- 27:48up.
- 27:50>> Okay. So, do you have any question for
- 27:52me? Like every I'm done from my side.
- 27:55>> Okay. So, thank you. Like usually I
- 27:57don't have any question, but just a
- 27:58small thing that I wanted to understand.
- 28:00Like since this is a fintech company, so
- 28:03I just I was a bit curious to know like
- 28:05how the security is managed from a
- 28:07engineering perspective.
- 28:09>> So, yeah, you were right. Like since we
- 28:11are fintech company, we have a, you
- 28:12know, a billion transaction happening in
- 28:15a month.
- 28:16So, security is a core part of our
- 28:19engineering culture. We basically here
- 28:21we follow a DevSecOps approach with
- 28:23shift-left security.
- 28:25We have a like automated security checks
- 28:27in CI/CD.
- 28:30Once those checks pass, then only code
- 28:32get promoted to next
- 28:35uh next environment. We have a strong
- 28:37access control across multiple account.
- 28:40We were using, you know, multiple cloud,
- 28:42like
- 28:43AWS, Azure, GCP for DR as well as uh
- 28:48for serving traffic to our clients. We
- 28:50have to Since we are, you know, fintech
- 28:52company, we have to
- 28:54uh follow multiple compliance here
- 28:57here here in our organization.
- 29:00And we have a like close collaboration
- 29:02between development, security, and
- 29:05operation team, and you will be part of
- 29:07that team as well.
- 29:09So, that's what we are doing here. Uh
- 29:12anything specific do you want to go?
- 29:14>> No, no, I think that's it. Apart from
- 29:16that, I just wanted to understand like
- 29:18are we into the phase of migration
- 29:19because I just uh got to know like from
- 29:21the question like mostly it was related
- 29:23to the migration. Are we using
- 29:25monolithic or something?
- 29:26>> yeah. So, we have a couple of migration
- 29:28going in our organization, and one of
- 29:30the migration is we are moving from ECS
- 29:33to EKS.
- 29:34So, that's why we are uh you know, I am
- 29:36asking question about migration
- 29:38especially uh Kubernetes.
- 29:42Because we want someone who worked with
- 29:45Kubernetes before extensively to help in
- 29:48migration. Okay? So, that is the reason.
- 29:51Yeah.
- 29:52>> Understood. Yeah, yeah, I think that's
- 29:54it. That is all from my side. Yeah.
- 29:56>> Okay. So,
- 29:57>> Okay, then. Thank you so much, Ravi.
- 29:58Thank you so much.
- 29:59Thank you so much. Bye. Yeah.
About this transcript
This page contains the full transcript of DevOps Live Interview - Fintech - 35 LPA | DevOps | Cloud | #interview #devops #aws #docker #cloud by jadeja rajpal sinh, generated from the public captions YouTube serves with the video. The transcript has 5,924 words across 975 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.