YouTube2Text

DevOps Live Interview - Fintech - 35 LPA | DevOps | Cloud | #interview #devops #aws #docker #cloud — Transcript

by jadeja rajpal sinh · 5,924 words · 975 segments · language en · Watch on YouTube

Full transcript

  1. 0:01Hello, Jadeja.
  2. 0:03>> Yeah, hey Ravi.
  3. 0:04>> Hey, how are you?
  4. 0:06>> I'm really good. How are you?
  5. 0:08>> I am also good. So, let me open your
  6. 0:11resume.
  7. 0:13Uh meanwhile, you can introduce
  8. 0:14yourself. Okay?
  9. 0:16>> Yeah. So, basically, my name is
  10. 0:18Jadeja Rajpasingh and I'm a DevOps
  11. 0:19engineer with around 7 year plus
  12. 0:20experience in designing, building, and
  13. 0:22managing all sort of highly available
  14. 0:24cloud platform.
  15. 0:25So, my core expertise lies in
  16. 0:27implementing the key pillars of DevOps,
  17. 0:28that is automation, continuous
  18. 0:30integration delivery, infrastructure as
  19. 0:32a code, monitoring, observability, and
  20. 0:34of course, the security part. So,
  21. 0:35throughout my career, I have worked on
  22. 0:37multiple cloud platform, but uh
  23. 0:39extensively on the side of AWS and a bit
  24. 0:41of GCP and Azure.
  25. 0:43So, I'm good at designing the cloud
  26. 0:44native platform. I worked on
  27. 0:46Kubernetes-based environment and set up
  28. 0:48the end-to-end uh CI/CD ecosystem for
  29. 0:50the enterprise application.
  30. 0:51I have successfully contributed to
  31. 0:53several large-scale transformation
  32. 0:55project like including the migration of
  33. 0:57Kubernetes cluster from GKE to Amazon
  34. 0:58EKS, GKE to EKS. I have modernized
  35. 1:01monolithic application from micro-based
  36. 1:04into the microservice-based
  37. 1:05architectures.
  38. 1:06Then, uh on the governance side, I have
  39. 1:08implemented the secure multi-account
  40. 1:10cloud architecture, centralized identity
  41. 1:12and access management, and solution
  42. 1:14aligned with the complex standards such
  43. 1:15as GDPR and HIPAA. And uh apart from it,
  44. 1:18I've also been involved into the client
  45. 1:19discussion and pre-sales uh engagement
  46. 1:21where I need to uh get in touch with my
  47. 1:23client and need to convert their
  48. 1:25business requirement into the technical
  49. 1:26architectures. And uh throughout this
  50. 1:29journey, I did some sort of
  51. 1:29certification like uh AWS Certified
  52. 1:31Solutions Architect Professional, Red
  53. 1:33Hat Certified System Administration, and
  54. 1:35Red Hat Certified Ansible Automation.
  55. 1:37And beyond all these technical
  56. 1:38responsibilities, I have trained more
  57. 1:40than uh 50 plus engineers in my uh
  58. 1:42current organization.
  59. 1:44>> Okay, yeah.
  60. 1:46So, uh you mentioned in your
  61. 1:48introduction that you have trained over
  62. 1:5050 DevOps engineer.
  63. 1:52So, could you tell me more about your
  64. 1:54role in that?
  65. 1:56>> So, like basically, like I've worked at
  66. 1:59a service-based company. So, every year
  67. 2:01around 20 or 30 DevOps engineer join the
  68. 2:03company.
  69. 2:04So, as a part of their onboarding, I
  70. 2:05conduct the technical sessions covering
  71. 2:07core DevOps concepts like the basic of
  72. 2:09AWS, Kubernetes, CICD, Terraform, and
  73. 2:11many more tools. So, my responsibility
  74. 2:13is to help them help the new engineer
  75. 2:15build a strong foundation so that they
  76. 2:16can contribute effectively to the client
  77. 2:19project. So, that is what usually we
  78. 2:21follow a procedure into our current
  79. 2:22organization.
  80. 2:23>> Okay, yeah, yeah. Okay.
  81. 2:25So, in your resume you mentioned that
  82. 2:27you worked in a lot of migration
  83. 2:29project, right?
  84. 2:31>> Mhm, yes, correct.
  85. 2:32>> So, yeah, let me give you a scenario.
  86. 2:34So, your company has decided to migrate
  87. 2:37the application from EC2 to
  88. 2:40Kubernetes.
  89. 2:41So, how will you decide Kubernetes is
  90. 2:43the right choice?
  91. 2:46>> Mhm, so like before even though going to
  92. 2:49the Kubernetes, firstly I will try to
  93. 2:51understand the current application and
  94. 2:52the problem that you are trying to
  95. 2:53solve.
  96. 2:54So, we need to check the application
  97. 2:55compatibility whether the application
  98. 2:57can be containerized easily or not
  99. 2:59without having the major changes. And
  100. 3:02like after that like what I will do,
  101. 3:03I'll compare the benefit of the
  102. 3:04Kubernetes like the auto scaling,
  103. 3:07self-healing, and the zero downtime
  104. 3:08deployment. But along with that, we need
  105. 3:09to also
  106. 3:10like keep one factor in mind that is the
  107. 3:12operational complexity and the cost
  108. 3:14because if like whether the improved
  109. 3:16scalability and the automation like it
  110. 3:18should justify the additional cost as
  111. 3:19well.
  112. 3:20Then what I will do, I'll just simply
  113. 3:23do a small POC with a non-production
  114. 3:25application. And if the results are
  115. 3:26positive, then I will create a proper
  116. 3:28document and recommend like with moving
  117. 3:30to the Kubernetes.
  118. 3:33>> Yeah, okay, yeah. Sounds good. So, now
  119. 3:36you decided Kubernetes is a right choice
  120. 3:38after doing POC and you know, creating a
  121. 3:40document. But like how would you plan
  122. 3:43the migration to ensure there is a no
  123. 3:45downtime for end user?
  124. 3:48>> So, firstly, the first step would be
  125. 3:50that we need to containerize the
  126. 3:51application and test it thoroughly. Then
  127. 3:54I will set up the Kubernetes
  128. 3:55infrastructure. It can be on anything,
  129. 3:56EKS,
  130. 3:58Azure AK AKS or GK or anything. Along
  131. 4:01with that, we need to properly set up
  132. 4:03the networking, ingress, storage and
  133. 4:05everything. Then what I will do I will
  134. 4:06just deploy the application in a
  135. 4:07non-production environment and perform
  136. 4:09some sort of functional and security
  137. 4:11testing over there. Then I will set up
  138. 4:12the CI/CD pipeline for automated
  139. 4:14deployments. Then at the end I will use
  140. 4:17a gradual migration strategy like the
  141. 4:18blue/green or the canary deployment
  142. 4:20instead of switching all the traffic at
  143. 4:21once. Then I will monitor the
  144. 4:23application health, the logs, some sort
  145. 4:24of matrices during the migration.
  146. 4:26And at the end like I will keep a
  147. 4:28rollback plan ready so that if any
  148. 4:29issues occur at that time while doing
  149. 4:31the migration, traffic can immediately
  150. 4:33be rolled back to the EC2 instance. So
  151. 4:35this is what usually I will
  152. 4:36follow a procedure to migrate the
  153. 4:39application.
  154. 4:41>> Okay, yeah. So let's assume after
  155. 4:43migration, deployment becomes slower
  156. 4:46than EC2.
  157. 4:48So how would you troubleshoot
  158. 4:50that?
  159. 4:52>> Okay, so after the migration, right? So
  160. 4:53firstly I will try to
  161. 4:55identify where the delay is happening
  162. 4:57because Kubernetes itself doesn't
  163. 4:59necessarily make the deployment slower.
  164. 5:00So it can be multiple reasons. So
  165. 5:01firstly I would check at the level of
  166. 5:03CI/CD because inside the CI/CD there we
  167. 5:06use multiple stages like the image
  168. 5:07build, image push, deployment. So
  169. 5:09firstly I will try to analyze if
  170. 5:10something is going wrong or maybe if
  171. 5:12something is taking bit more time into
  172. 5:14the form of CI/CD. That is my first
  173. 5:16thing that is that is my first thing
  174. 5:18that I will look into it. Then I will
  175. 5:19verify if the images are sizes are too
  176. 5:21large, then I will try to optimize the
  177. 5:22Docker file using the multi-stage build
  178. 5:24and smaller base images. So that will
  179. 5:26ultimately help to reduce the deployment
  180. 5:28time. Then the other factor that can
  181. 5:31also cause the slow deployment into the
  182. 5:33form of Kubernetes that can be like the
  183. 5:35the pod scheduling. Some sort of
  184. 5:37sometimes the pod remains in the pending
  185. 5:39pending state because of insufficient
  186. 5:41CPU or memory or the node availability.
  187. 5:43Then inside the Kubernetes I will try to
  188. 5:45review the readiness and liveliness
  189. 5:47probe as well because incorrect probe
  190. 5:49configuration can also increase the
  191. 5:51delay time.
  192. 5:52Then at the end, I will try to monitor
  193. 5:54the cluster resource like I can enable
  194. 5:55some sort of cluster auto scaler so that
  195. 5:57if any nodes are going to the bottleneck
  196. 5:58or if any nodes is creating an issue, it
  197. 6:01will automatically get scaled up. So,
  198. 6:03usually this would be some of my
  199. 6:04checklist here.
  200. 6:06>> Okay, yeah. Okay. So, now during a
  201. 6:09deployment,
  202. 6:11all old pod were terminated before new
  203. 6:13pod become ready.
  204. 6:15So, that is causing down time. So, how
  205. 6:18would you prevent this?
  206. 6:20So, it's not happened in future.
  207. 6:22>> So, even though in in many of our
  208. 6:24project like we have faced the same
  209. 6:25issue. So, usually this happens due to a
  210. 6:27incorrect deployment strategy or you can
  211. 6:29say that some sort of incorrect
  212. 6:31configuration inside the Kubernetes. So,
  213. 6:33if you want to if you wanted to prevent
  214. 6:34such type of situation, then the first
  215. 6:36thing would be like we need to
  216. 6:38properly set up the readiness probes
  217. 6:39inside the Kubernetes so that only the
  218. 6:41traffic will reach to the healthy pods.
  219. 6:43And also like I will
  220. 6:45when I will use the rolling update
  221. 6:46strategy, so there we need to also
  222. 6:47mention the max available and the max
  223. 6:49surge value so that the pod will not get
  224. 6:51terminated until the new pods are
  225. 6:53getting ready over there.
  226. 6:54And there is one more policy that we
  227. 6:56could set up inside the Kubernetes
  228. 6:57cluster that is the pod disruption
  229. 6:59budget so that it maintains a minimum
  230. 7:01number of available pods during the
  231. 7:02deployment. And lastly like for the
  232. 7:05critical application deployment, we can
  233. 7:07go ahead with the blue green or the
  234. 7:09canary deployment like I usually I
  235. 7:10prefer the canary deployment to achieve
  236. 7:12the zero down time deployment.
  237. 7:14>> Yeah, it will avoid the down time. Yeah,
  238. 7:16okay. So, now you know, there is a
  239. 7:18situation where some application should
  240. 7:21not run together on the same node due to
  241. 7:24XYZ reason. So,
  242. 7:27so because uh
  243. 7:29sometime they consume too much CPU
  244. 7:32or
  245. 7:33So, how would you ensure proper
  246. 7:35placement across node?
  247. 7:39>> So, like like like, say if there are
  248. 7:41five pods for a same application or
  249. 7:43different type of application and we
  250. 7:44want each application to go into the
  251. 7:46different nodes, right?
  252. 7:47So, also, we can also put some sort of
  253. 7:49condition that no two pods get together
  254. 7:51into the pods, so we can use the pod
  255. 7:52anti-affinity over there. So, like, by
  256. 7:54doing this, like, our
  257. 7:56like, it will automatically,
  258. 7:58like,
  259. 7:59it will it will like, only if if one pod
  260. 8:01is going into one of the node, then pod
  261. 8:03anti-affinity will give us the option or
  262. 8:05the feasibility that no other pod would
  263. 8:06get scheduled into that specific node
  264. 8:08over over there.
  265. 8:09And if the application needs to run only
  266. 8:11one only on a specific node, then over
  267. 8:12there I'll label the nodes and then we
  268. 8:14can use the node selector or the node
  269. 8:16affinity over there into our
  270. 8:17configuration file. So, this is how we
  271. 8:19can do that.
  272. 8:21>> Okay. So, after migration, the this is a
  273. 8:24scenario like developer is requesting
  274. 8:27shell access to production pod for
  275. 8:29debugging.
  276. 8:30So, would you allow it?
  277. 8:32>> Oh, no, no, actually not at all because
  278. 8:34I would not allow any sort of direct uh,
  279. 8:36shell access to the production pod by
  280. 8:37default because it will introduce the
  281. 8:39security and the compliances.
  282. 8:40So, like, what I would encourage, like,
  283. 8:43I would encourage the developer to use
  284. 8:44the logs, to use the metrics or maybe
  285. 8:47whatever the tool that we are using,
  286. 8:48Prometheus, Grafana, DataDog, whatever
  287. 8:49the tools that we are using. So, maybe
  288. 8:51they can check the logs from there.
  289. 8:53But if still something is there, if the
  290. 8:55developer like, insisting me to give the
  291. 8:57permission if they really need it, so I
  292. 8:59will
  293. 9:00provide the temporary control access
  294. 9:01using the RBA with proper approval like
  295. 9:03from my management and everyone. Then
  296. 9:05also, along with the access, I will try
  297. 9:07to log and audit all the actions what
  298. 9:08they're performing with that specific
  299. 9:10access. [laughter]
  300. 9:11And after like, once the debugging is
  301. 9:13completed, then I will immediately
  302. 9:14revoke those temporary access over
  303. 9:16there.
  304. 9:17>> Okay, okay. Yeah.
  305. 9:18So, now now imagine your Docker registry
  306. 9:21become unavailable during a one of your
  307. 9:24production deployment.
  308. 9:26So, what impact would you expect after
  309. 9:28this incident?
  310. 9:31>> Okay, so image registry is going down.
  311. 9:32So, firstly, I'll check whether the
  312. 9:33required Docker image is already present
  313. 9:35on the worker node. And if the image is
  314. 9:37already there, the existing code will
  315. 9:39still run without any sort of impact.
  316. 9:41But if the Kubernetes is trying to
  317. 9:43deploy a new version, so the deployment
  318. 9:44will fail and the new code will not get
  319. 9:46started over there. So to minimize the
  320. 9:48risk, what I can do, I can simply use
  321. 9:51some sort of highly available or the
  322. 9:52private registry such as ECR so that we
  323. 9:54can also like at the same time we can
  324. 9:56also enable some sort of image caching
  325. 9:57where possible and we can also perform
  326. 9:59some sort of disaster recovery plan for
  327. 10:01the image
  328. 10:04uh that image ratio too so that if the
  329. 10:06primary region goes down,
  330. 10:08there would be a backup
  331. 10:09registry in another region so that we
  332. 10:11could we can pull the images from there.
  333. 10:13So that is what we can do easily.
  334. 10:15>> Okay, yeah.
  335. 10:16So now assume your one of your teammate
  336. 10:19accidentally deleted
  337. 10:21the production name space.
  338. 10:23Uh
  339. 10:24what will you do? Like Kubernetes name
  340. 10:26space. So in that case, what will you
  341. 10:28do?
  342. 10:30>> Okay, so they deleted the production
  343. 10:31name space only, right? So basically
  344. 10:32firstly I will stop the ongoing
  345. 10:34deployments or the changes to prevent
  346. 10:36the further issues.
  347. 10:37Then I will check if there is any sort
  348. 10:38of backup tool like generally we use the
  349. 10:40Valero in our Kubernetes cluster so that
  350. 10:42it takes the backup of everything like
  351. 10:44it takes the backups of name space,
  352. 10:45resource and everything. But let's say
  353. 10:47we don't have a backup. So if there's no
  354. 10:49backup exists, then I will simply
  355. 10:50redeploy the name space and workload
  356. 10:52using the GitHub repository because
  357. 10:53since we are deploying everything with
  358. 10:54the help of Argo CD. So we are having
  359. 10:56all the manifest file, we are having all
  360. 10:58the answer of each and every
  361. 10:59microservice. And then I will verify
  362. 11:01that all the applications are healthy
  363. 11:02and working like as expected after the
  364. 11:04restore.
  365. 11:05And then finally, I will try to
  366. 11:07investigate the root cause and implement
  367. 11:08some sort of
  368. 11:09safeguard like Rback, some sort of
  369. 11:11approval workflow and some sort of
  370. 11:13backup policy to prevent the similar
  371. 11:14incident so that no one would be able to
  372. 11:16delete the production name space again.
  373. 11:18>> Okay, okay. So let's say let's assume
  374. 11:21after migration is done, after 60 year
  375. 11:24after 6 month, management ask whether
  376. 11:27Kubernetes has delivered value. So in In
  377. 11:30case, what KPI would you present to
  378. 11:32management?
  379. 11:34>> Okay, so we wanted to compare whether
  380. 11:35the migration that we did that is
  381. 11:37successful or not, right?
  382. 11:38>> Yeah, right.
  383. 11:39>> So,
  384. 11:40there are multiple factors over there.
  385. 11:41So, the first thing that would be the
  386. 11:43deployment time. Like I will try to
  387. 11:44compare the before and the after the
  388. 11:46migration of the deployment time, like
  389. 11:47which one is taking the lesser
  390. 11:48deployment time. Then I will also try to
  391. 11:50check some sort of application
  392. 11:51availability, some sort of resource
  393. 11:53utilization, like when we are going with
  394. 11:54the Kubernetes, so whether it is
  395. 11:56optimizing the resources or not. Then
  396. 11:58the major like one of the major factor
  397. 11:59that is the infrastructure cost. Like I
  398. 12:02will try to compare the cost before and
  399. 12:04after the migration. Then also we need
  400. 12:06to check like how like sometimes the
  401. 12:07scale like sometimes there might be a
  402. 12:09chance that the costing has been
  403. 12:10increased because we are implementing
  404. 12:11some sort of application scalability
  405. 12:12feature as well. So, we need to also
  406. 12:14understand like how quickly that it
  407. 12:16handles the traffic spike. So, it
  408. 12:17depends on the business requirement as
  409. 12:18well. Then at the end I will try to like
  410. 12:21some I will try to check some sort of
  411. 12:23operational matrices like the deployment
  412. 12:24success rate, some sort of
  413. 12:26rollback frequency, some sort of MTTR,
  414. 12:29which is the mean time recovery. So,
  415. 12:31that are some of my KPIs that I would
  416. 12:33check before
  417. 12:34giving the conclusion to them.
  418. 12:38>> Okay.
  419. 12:39So, in your CI/CD, your appli- it is
  420. 12:42showing your application is deployed
  421. 12:44successfully. But for some user
  422. 12:47still see the old version.
  423. 12:50So, in that case, what could be the
  424. 12:52reason why some of some users are still
  425. 12:55seeing the old version of application?
  426. 12:58>> let's say if if you're working into the
  427. 13:00production environment and if the
  428. 13:01infrastructure is quite big. So, firstly
  429. 13:02I will check whether the deployment is
  430. 13:04like whenever we make the deployment, so
  431. 13:06sometimes it takes the time to roll out
  432. 13:08all the deployments all the ports. So, I
  433. 13:10will check whether all the ports are
  434. 13:11still in the progress or if all the
  435. 13:13ports are has been updated. And let's
  436. 13:15let's say if everything is updated, then
  437. 13:16I will verify the service like whether
  438. 13:17it is routing the traffic correctly to
  439. 13:19the new port or not.
  440. 13:20Then at the same time we need to check
  441. 13:22some sort of things into the side of
  442. 13:25browser, some sort of CDN, some sort of
  443. 13:26caching. If you're any sort of caching,
  444. 13:28then maybe we need to invalidate the
  445. 13:30caches so that it
  446. 13:32serve the new content.
  447. 13:34And one more thing, like if we are using
  448. 13:36the canary or the blue/green deployment,
  449. 13:38so it's expected that the some users see
  450. 13:40the old version and while others see the
  451. 13:42new version until that specific
  452. 13:44deployment goes complete or like whether
  453. 13:45that specific deployment get completed.
  454. 13:48So sometimes they like we can
  455. 13:50face some sort of issues over there.
  456. 13:53>> Mhm, okay. That's yeah, correct. So now
  457. 13:56how would you design your CI/CD pipeline
  458. 13:58so bad deployment never reaches to
  459. 14:00production directly?
  460. 14:03>> Okay. So like firstly I would say that
  461. 14:05we cannot guarantee that a faulty
  462. 14:07deployment will never reach the
  463. 14:08production.
  464. 14:09But yeah, we can reduce the chances by
  465. 14:11adding some sort of proper validation
  466. 14:12step inside the pipeline. So inside the
  467. 14:14staging like we can integrate some sort
  468. 14:16of unit test cases, some sort of
  469. 14:17security test cases inside the pipeline.
  470. 14:19Then before pushing the images we could
  471. 14:21scan the images so that if any
  472. 14:23vulnerability is there inside that
  473. 14:24image, we could get to know before
  474. 14:26making a deployment.
  475. 14:27Uh then
  476. 14:28like also I will not deploy any any sort
  477. 14:30of things or any deployment directly to
  478. 14:32the production. Firstly I will try to
  479. 14:33validate into the lower environment that
  480. 14:35is development or the staging
  481. 14:36environment. And like if required we can
  482. 14:39also integrate some sort of manual
  483. 14:41approval inside the production
  484. 14:42deployment as well.
  485. 14:43Then at the same time as already
  486. 14:45mentioned that we can go ahead with some
  487. 14:46sort of canary and the blue/green
  488. 14:47deployment so that it will minimize the
  489. 14:49risk. It won't commit like we can't give
  490. 14:51the commitment that the faulty
  491. 14:53deployment will never reach but at least
  492. 14:54we could minimize the risk by doing all
  493. 14:56such status of things.
  494. 14:58>> Yes, you have to add multiple gates on
  495. 15:00CI, right?
  496. 15:01>> Yep, correct.
  497. 15:02>> So now you know there you have multiple
  498. 15:04micro service in your application. Like
  499. 15:06it could be 50 or
  500. 15:0850 plus, let's assume.
  501. 15:10So for that would you create a separate
  502. 15:13CI/CD pipeline for each one?
  503. 15:16>> Um like
  504. 15:18to be honest, it depends on the project.
  505. 15:19But in most cases, yes, like I create a
  506. 15:22separate pipeline for each microservice.
  507. 15:23So, this allows a team to develop, test,
  508. 15:26and deploy their independent services
  509. 15:27without affecting the other one. But,
  510. 15:29let's say I will not duplicate the
  511. 15:30pipeline logic. I will create a reusable
  512. 15:32pipeline template, and so that each
  513. 15:34service can use the same template with a
  514. 15:35different configuration. So, with the
  515. 15:38help of this particular reusable
  516. 15:40pipeline template, it will keep the
  517. 15:41pipeline consistent and easier to
  518. 15:42maintain. And it follows a dry
  519. 15:44principle, which is don't repeat
  520. 15:45yourself. So, I think so, yeah, this is
  521. 15:47usually what I do.
  522. 15:49>> Mhm, separate pipeline because also, you
  523. 15:52know, rollback become easier, right?
  524. 15:55>> Correct. But, at the same time, if if we
  525. 15:57are using any sort of reusable
  526. 15:58component,
  527. 15:59then I will create a specific template
  528. 16:01for it, so that they can use those
  529. 16:03reusable components from that specific
  530. 16:04template.
  531. 16:06>> Okay, okay, yeah.
  532. 16:08So, now, like, you know, your company
  533. 16:10decided to migrate pipelines from
  534. 16:13Jenkins to GitHub Action due to XYZ
  535. 16:16reason, due to license expire, or
  536. 16:18whatever. So, how do you ensure a smooth
  537. 16:21migration without affecting existing
  538. 16:23deployments?
  539. 16:25>> So, firstly, we need to understand the
  540. 16:28existing Jenkins pipeline, the
  541. 16:29integration, the deployment process, the
  542. 16:31plugin, the users, however they're set
  543. 16:34up. So, firstly, I will try to
  544. 16:35understand the existing platform.
  545. 16:36Then, I will try to set up the same
  546. 16:38thing into the GitHub Action, then I
  547. 16:39will migrate one non-critical
  548. 16:40application
  549. 16:41just for a POC purpose. And then, I will
  550. 16:43run the Jenkins and the GitHub Action in
  551. 16:44parallel for both of like for some time,
  552. 16:46so that we could compare the result. And
  553. 16:48once the new pipeline is stable, then I
  554. 16:49will gradually migrate the remaining
  555. 16:51applications over there. And throughout
  556. 16:53the migration, I will closely monitor
  557. 16:55our deployment like and keep the Jenkins
  558. 16:56as a rollback option until the migration
  559. 16:58is fully completed. So, usually, this is
  560. 17:00what I will do.
  561. 17:01>> Okay.
  562. 17:03So,
  563. 17:04there is a scenario, okay? So, your
  564. 17:06application must be deployed
  565. 17:08simultaneously to
  566. 17:10multi-cloud, like, let's assume AWS and
  567. 17:13Azure.
  568. 17:14So, Uh, one deployment succeeded, but
  569. 17:17the other fails. So, how would your uh
  570. 17:20pipeline handle this scenario?
  571. 17:23>> Um, so firstly I will just stop the
  572. 17:25pipeline and identify where the why the
  573. 17:27Azure deployment getting failed. Then I
  574. 17:29will not mark the deployment as
  575. 17:30successful until both AWS and the Azure
  576. 17:33deployment gets
  577. 17:34completed successfully. And if they need
  578. 17:36if let's say if AWS is already running
  579. 17:38the new version, so firstly I will like
  580. 17:40I will either roll back it to the
  581. 17:42previous one or
  582. 17:44uh, like we depend like on the business
  583. 17:45requirement. So, if we can roll back it
  584. 17:47to the previous version, we can simply
  585. 17:48do it or else I will just keep it keep
  586. 17:50it as as it is running. And then I will
  587. 17:52try to fix the Azure issue and then
  588. 17:54rerun the pipeline and ensure that both
  589. 17:55environment are on the same new
  590. 17:57application version. And then at the end
  591. 17:59like I will try to monitor some sort of
  592. 18:00both deployment to verify whether they
  593. 18:02are healthy before completing the
  594. 18:03release. So, like it it will be a part
  595. 18:06of CICD pipeline. Like firstly the AWS
  596. 18:08pipeline will get triggered. Sorry, AWS
  597. 18:09pipeline will get deployed successfully
  598. 18:11then the Azure will get deployed. And if
  599. 18:13if both the deployments are getting
  600. 18:15successful, then I will make the
  601. 18:16deployment as successful otherwise I
  602. 18:17will just roll back for both of them.
  603. 18:19So, this is what we can do.
  604. 18:22>> Okay. So, do you have experience with
  605. 18:24Terraform as well or IAC tool?
  606. 18:26>> yeah. So, even though like for like from
  607. 18:28last 3-4 years I'm extensively working
  608. 18:30into the side of Terraform only whatever
  609. 18:31the infrastructure that I'm creating it
  610. 18:32up.
  611. 18:33>> Okay.
  612. 18:33>> Yeah.
  613. 18:34>> So, now your organization organization
  614. 18:36decided to manage all their
  615. 18:38infrastructure using Terraform only. But
  616. 18:41most of the resources were created
  617. 18:43manually. So, how would you approach the
  618. 18:46migration? Like yeah.
  619. 18:48>> Understood. So, the the resources are
  620. 18:50already there into the side of AWS,
  621. 18:51right? So,
  622. 18:52>> AWS. Yeah.
  623. 18:53>> Yeah. So, firstly I will try to identify
  624. 18:54and prioritize the existing AWS
  625. 18:56resources. I will try to I will write
  626. 18:57the Terraform code for those resources.
  627. 18:59And instead of recreating them, I will
  628. 19:01use the Terraform import.
  629. 19:04With help of import Terraform import you
  630. 19:05you can bring the existing resource into
  631. 19:07your Terraform state file. Then I will
  632. 19:09try to run the Terraform plan to verify
  633. 19:10that the Terraform doesn't propose any
  634. 19:12sort of unexpected changes. And once
  635. 19:14everything is validated, I will simply
  636. 19:15uh use the Terraform as a single source
  637. 19:17of truth for all the future
  638. 19:18infrastructure changes. So, we need to
  639. 19:20go into the repetition. Like like we
  640. 19:22will make the changes. Then again, we
  641. 19:23will run the Terraform plan. We need to
  642. 19:24go into the loop until and unless it
  643. 19:26doesn't show us any sort of
  644. 19:28uh
  645. 19:28uh changes over there. So, we need to
  646. 19:30repeat those changes again and again.
  647. 19:32>> Yeah, you have to run the Terraform
  648. 19:34import. Then you have to add the code on
  649. 19:36Terraform file.
  650. 19:38And then
  651. 19:38>> Correct.
  652. 19:39>> Repeat.
  653. 19:39>> to run the Terraform plan.
  654. 19:41And ultimately, I think it should share
  655. 19:43it should uh share share that like there
  656. 19:45is no infrastructure changes. Then only
  657. 19:47we would be able to 100% sure that it's
  658. 19:49working fine.
  659. 19:51>> Okay. Okay. So, now uh how do you
  660. 19:53prevent multiple engineers from making
  661. 19:55Terraform changes to the same
  662. 19:57environment simultaneously?
  663. 19:59>> So, it's a default feature of the
  664. 20:00Terraform. So, we can simply use the
  665. 20:02remote backend such as a S3 bucket to
  666. 20:03store the Terraform state file. And I
  667. 20:05would simply enable the state locking
  668. 20:06using the using the DynamoDB so that
  669. 20:08only one Terraform operation can run at
  670. 20:10a specific time. And also, I will ensure
  671. 20:13that all the Terraform changes goes
  672. 20:14through the CI/CD pipeline. Instead of
  673. 20:16being running it from the manually. So,
  674. 20:18inside the CI/CD pipeline, if someone
  675. 20:19has triggered something, then the second
  676. 20:21pipeline would automatically go into the
  677. 20:23queue. And also, at the same time, I'll
  678. 20:24separate the workspace or the state file
  679. 20:25by a for different different environment
  680. 20:27like the dev, QA, production. And also,
  681. 20:29like if required, we could also
  682. 20:30mention the code review and approval
  683. 20:32before applying the infrastructure
  684. 20:34changes over there.
  685. 20:36>> Okay. Yeah, so in multi-cloud scenario,
  686. 20:39so now uh like uh you need to provision
  687. 20:41infrastructure across multiple
  688. 20:43multi-cloud like AWS, Azure, GCP using
  689. 20:46Terraform only, okay?
  690. 20:48So, how would you organize the code base
  691. 20:50to maximize reusability or
  692. 20:52maintainability?
  693. 20:55>> So, firstly, I will just separate the
  694. 20:57code by like I will separate the code by
  695. 20:59cloud provider like the AWS, Azure, GCP
  696. 21:01while keeping a common repository
  697. 21:03structure over there. I will try to
  698. 21:04create a reusable Terraform modules over
  699. 21:06there so that for common services like
  700. 21:08the
  701. 21:09uh networking, some sort of security,
  702. 21:10some sort of computer storage, we could
  703. 21:12use the those usable Terraform modules.
  704. 21:15And then I will keep the environment
  705. 21:16specific configuration like the dev, QA,
  706. 21:17prod for using the different type of
  707. 21:19variable files.
  708. 21:21Then I will maintain some sort of
  709. 21:22separate state file for each cloud and
  710. 21:24each environment to avoid the conflict.
  711. 21:25Then like we already discussed like I
  712. 21:27will use the remote backend with the
  713. 21:29state locking and manage all the
  714. 21:30deployment with the help of CICD
  715. 21:31pipeline. So this is what I could do.
  716. 21:35>> Okay, yeah. With state lock locking and
  717. 21:37everything. Yeah, okay.
  718. 21:38So now uh
  719. 21:40how would you prevent a developer from
  720. 21:42accidentally deleting production
  721. 21:44resources using Terraform? Because you
  722. 21:46know, a Terraform code will now store on
  723. 21:49GitHub or version control system. So how
  724. 21:51would you prevent it? Yeah.
  725. 21:53>> So firstly, like we need to follow the
  726. 21:55least privilege principle. So I will
  727. 21:57restrict the production access using the
  728. 21:58IM role.
  729. 21:59Then along with that, I will also ensure
  730. 22:01that all Terraform changes goes through
  731. 22:03a CICD pipeline and mandatory code
  732. 22:05review and approval.
  733. 22:06Then inside the Terraform, we could use
  734. 22:08some sort of block like the prevent
  735. 22:10destroy is equal to two true. So that it
  736. 22:12will not let us delete the critical
  737. 22:14resource like the database, some sort of
  738. 22:16production cluster.
  739. 22:17Then I will also require developer to
  740. 22:18review the Terraform plan before they
  741. 22:21apply any sort of before they apply any
  742. 22:22sort of changes over there. And yeah, I
  743. 22:24think so this is And also we could
  744. 22:25separate out the production and
  745. 22:26non-production environment for the
  746. 22:28better practice. Yeah.
  747. 22:29>> You have You have to add life cycle
  748. 22:31policy.
  749. 22:32>> Life cycle policy. That would be the
  750. 22:33prevent destroy It's the life cycle
  751. 22:35policy over there. Yeah.
  752. 22:36>> Mhm.
  753. 22:37Get it. So now you know, have you worked
  754. 22:40with FinOps as well?
  755. 22:42Implemented FinOps in your
  756. 22:44organization before?
  757. 22:45>> Like whenever Yeah, usually we follow
  758. 22:47the like we worked into the side of
  759. 22:48FinOps as well where we need to automate
  760. 22:49the infrastructure cost and everything
  761. 22:51for our client.
  762. 22:53>> So now okay, let me give you a scenario.
  763. 22:54So now your company want to reduce cost
  764. 22:56by 40%. Let us assume AWS cloud, okay?
  765. 23:01>> Okay.
  766. 23:01>> But they don't want to they don't want
  767. 23:03any impact on application performance.
  768. 23:06So where would you start?
  769. 23:09>> So like I will analyze the current AWS
  770. 23:12cost using the cost explorer and
  771. 23:14identify the top cost consuming
  772. 23:16services.
  773. 23:17Then I would check whether the EC2
  774. 23:18instance EKS nodes RDS or storage are
  775. 23:21over provisioned and try to right size
  776. 23:23them. Okay.
  777. 23:25Then we need to also check like we we
  778. 23:26can improve the resource utilization by
  779. 23:28setting up some sort of proper condition
  780. 23:30like the CPU memory some sort of request
  781. 23:32limit into and we can also try to remove
  782. 23:34the ideal resources. If we are not using
  783. 23:36them it should not be there into our
  784. 23:37cloud. Then I will try to use some sort
  785. 23:39of auto scaling some sort of cluster
  786. 23:40auto scaling also if you wanted some
  787. 23:43sort of advanced
  788. 23:45scaling then we can use carpenter to
  789. 23:47ensure that only the resource that we
  790. 23:48actually need so with the help of
  791. 23:50carpenter provides us some sort of
  792. 23:51advanced feature over there. And for
  793. 23:53some sort of non-critical workloads like
  794. 23:55the dev QL Q Q environments so I we
  795. 23:58could also try to implement some sort of
  796. 24:00sport instance machine over there. And
  797. 24:02also we could go ahead with some sort of
  798. 24:03reserve plan and the saving plan as well
  799. 24:05because this is also one of the
  800. 24:07factor that we could consider and like I
  801. 24:09would continuously monitor some sort of
  802. 24:11cost and like utilize utilization to
  803. 24:13ensure that the optimization doesn't
  804. 24:15affect the application performance.
  805. 24:17>> Okay. Yeah.
  806. 24:20So now
  807. 24:21you know in one of your AWS account your
  808. 24:24primary AWS region become suddenly
  809. 24:26unavailable due to war or whatever.
  810. 24:30So how would your application recover in
  811. 24:33on that case?
  812. 24:34>> So even though right now I'm working for
  813. 24:36them like some of my clients and they
  814. 24:38are from US and Europe so over there we
  815. 24:39need to perform some sort of DR over
  816. 24:41there. So I will ensure that the
  817. 24:43disaster recovery setup is already there
  818. 24:44in place with a secondary AWS region. So
  819. 24:47if first region fails then the traffic
  820. 24:48will
  821. 24:49shift to the secondary region using the
  822. 24:50route 53 or there is another service
  823. 24:52that you can use that is the AWS global
  824. 24:54accelerator service. Then I will verify
  825. 24:56that the application database all all
  826. 24:58the other dependencies are healthy in
  827. 24:59the DR region. And if required, we could
  828. 25:01also restore the latest data from the
  829. 25:03backup or use the cross region
  830. 25:05replication for minimal data loss. And
  831. 25:07once the primary region is available
  832. 25:08again, then I will validate it and plan
  833. 25:11a controlled rollback over there.
  834. 25:12So, that is what we could do.
  835. 25:15>> Okay. So, in now in case of production
  836. 25:18alert, like severity one or P1 alert,
  837. 25:21so walk me through your first 30
  838. 25:24minutes. What action would you take on
  839. 25:26that case? Will you inform your manager
  840. 25:28or whatever action
  841. 25:30>> So, usually even though right now in our
  842. 25:32current
  843. 25:33>> project, usually we get an alert like
  844. 25:35there are some sort of SEV1, SEV2. That
  845. 25:37depends on the criticality criticality
  846. 25:39of that particular alert. So, we usually
  847. 25:41once I will get some sort of mail or
  848. 25:44some sort of like we usually have a
  849. 25:45portal where we would be able to see
  850. 25:47what are the sort of severity issues are
  851. 25:49happening in the production environment.
  852. 25:51So, firstly I will acknowledge the alert
  853. 25:52and assess the impact. Like is it I'll
  854. 25:55identify which application or the users
  855. 25:57are all the services are getting
  856. 25:58affected over there.
  857. 25:59Then I will simply join the incident
  858. 26:01bridge and inform the relevant team like
  859. 26:02the developer, infrastructure, the
  860. 26:04database team, whoever is involved over
  861. 26:05there. Then we'll check some sort of
  862. 26:07dashboard, some sort of logs, matrices,
  863. 26:09then try to check if any recent
  864. 26:11deployment to identify the possible root
  865. 26:13causes. And if a recent deployment
  866. 26:15causes the issue, then I will simply
  867. 26:17roll back it to the previous version to
  868. 26:19restore the services. And once the
  869. 26:20application is stable, I will continue
  870. 26:22investigating the root cause and monitor
  871. 26:24the system to ensure that the issue
  872. 26:26doesn't
  873. 26:27reoccur once again. And after the
  874. 26:29incident is resolved, I will conduct a
  875. 26:31root cause analysis that is RCA and
  876. 26:33implement the preventive measure on top
  877. 26:35of it.
  878. 26:37>> Okay. Yeah. Okay, so this is a last
  879. 26:39question from my side. Finally, I know
  880. 26:43finally if you decided to join our
  881. 26:45company as a senior DevOps engineer and
  882. 26:48you were given complete ownership of our
  883. 26:50platform,
  884. 26:52so what are the first three improvement
  885. 26:54you would look for before making any
  886. 26:57changes?
  887. 26:59>> Any changes, right? So, firstly, to be
  888. 27:01honest, like I will try to understand
  889. 27:03the current platform because I will also
  890. 27:05need some sort of time to review the
  891. 27:07architecture, CCD pipeline, the
  892. 27:09monitoring, security, infrastructure,
  893. 27:11infrastructure, and everything.
  894. 27:13Like there are many factors. Like you
  895. 27:14guys would be using monitoring,
  896. 27:16alerting, backup, disaster recovery
  897. 27:17plan. So, firstly, I will try to
  898. 27:19identify the current pain points over
  899. 27:21there and try to identify some sort of
  900. 27:23technical debt which is there in two of
  901. 27:24your existing platform.
  902. 27:27Then I will try to look for some sort of
  903. 27:29automation gap, some sort of cost
  904. 27:31optimization, some sort of performance
  905. 27:32bottleneck, and the developer experience
  906. 27:34improvement. I will try to check the I
  907. 27:36will try to take a feedback from the
  908. 27:37developer as well. Like are they happy
  909. 27:39with our system or not? Then I will
  910. 27:41prioritize changes based on the business
  911. 27:43impact instead of making the unnecessary
  912. 27:45modification. So, it usually depends
  913. 27:46like how the infrastructure has been set
  914. 27:48up.
  915. 27:50>> Okay. So, do you have any question for
  916. 27:52me? Like every I'm done from my side.
  917. 27:55>> Okay. So, thank you. Like usually I
  918. 27:57don't have any question, but just a
  919. 27:58small thing that I wanted to understand.
  920. 28:00Like since this is a fintech company, so
  921. 28:03I just I was a bit curious to know like
  922. 28:05how the security is managed from a
  923. 28:07engineering perspective.
  924. 28:09>> So, yeah, you were right. Like since we
  925. 28:11are fintech company, we have a, you
  926. 28:12know, a billion transaction happening in
  927. 28:15a month.
  928. 28:16So, security is a core part of our
  929. 28:19engineering culture. We basically here
  930. 28:21we follow a DevSecOps approach with
  931. 28:23shift-left security.
  932. 28:25We have a like automated security checks
  933. 28:27in CI/CD.
  934. 28:30Once those checks pass, then only code
  935. 28:32get promoted to next
  936. 28:35uh next environment. We have a strong
  937. 28:37access control across multiple account.
  938. 28:40We were using, you know, multiple cloud,
  939. 28:42like
  940. 28:43AWS, Azure, GCP for DR as well as uh
  941. 28:48for serving traffic to our clients. We
  942. 28:50have to Since we are, you know, fintech
  943. 28:52company, we have to
  944. 28:54uh follow multiple compliance here
  945. 28:57here here in our organization.
  946. 29:00And we have a like close collaboration
  947. 29:02between development, security, and
  948. 29:05operation team, and you will be part of
  949. 29:07that team as well.
  950. 29:09So, that's what we are doing here. Uh
  951. 29:12anything specific do you want to go?
  952. 29:14>> No, no, I think that's it. Apart from
  953. 29:16that, I just wanted to understand like
  954. 29:18are we into the phase of migration
  955. 29:19because I just uh got to know like from
  956. 29:21the question like mostly it was related
  957. 29:23to the migration. Are we using
  958. 29:25monolithic or something?
  959. 29:26>> yeah. So, we have a couple of migration
  960. 29:28going in our organization, and one of
  961. 29:30the migration is we are moving from ECS
  962. 29:33to EKS.
  963. 29:34So, that's why we are uh you know, I am
  964. 29:36asking question about migration
  965. 29:38especially uh Kubernetes.
  966. 29:42Because we want someone who worked with
  967. 29:45Kubernetes before extensively to help in
  968. 29:48migration. Okay? So, that is the reason.
  969. 29:51Yeah.
  970. 29:52>> Understood. Yeah, yeah, I think that's
  971. 29:54it. That is all from my side. Yeah.
  972. 29:56>> Okay. So,
  973. 29:57>> Okay, then. Thank you so much, Ravi.
  974. 29:58Thank you so much.
  975. 29:59Thank you so much. Bye. Yeah.

About this transcript

This page contains the full transcript of DevOps Live Interview - Fintech - 35 LPA | DevOps | Cloud | #interview #devops #aws #docker #cloud by jadeja rajpal sinh, generated from the public captions YouTube serves with the video. The transcript has 5,924 words across 975 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.