YouTube2Text

DevOps Live Interview | DevOps | Cloud | DevSecOps | SRE #interview #devops #aws #docker #cloud — Transcript

by jadeja rajpal sinh · 4,957 words · 817 segments · language en · Watch on YouTube

Full transcript

  1. 0:04Hi Jadeja.
  2. 0:05>> Yeah, Akash.
  3. 0:08>> So, let me open your resume once. Just
  4. 0:10give me a sec.
  5. 0:13Okay.
  6. 0:15So,
  7. 0:16till then till then please start by
  8. 0:18introducing yourself.
  9. 0:20>> Yeah. So, my name is Jadeja Rajpalsinh
  10. 0:22and I'm a DevOps engineer with 7 year
  11. 0:24plus experience in designing, building,
  12. 0:26operating scalable and highly available
  13. 0:28cloud platform.
  14. 0:29My core experience lies in implementing
  15. 0:31the key pillars of DevOps that is
  16. 0:33automation, continuous integration
  17. 0:34delivery, infrastructure as a code,
  18. 0:36monitoring, and security.
  19. 0:38Throughout my journey, I have worked on
  20. 0:40multiple cloud platform, but most
  21. 0:43extensively in the side of AWS and a bit
  22. 0:45of GCP and Azure. I have worked on
  23. 0:47designing cloud native platform,
  24. 0:49Kubernetes based environment, and
  25. 0:51end-to-end CCD ecosystem for enterprise
  26. 0:53application.
  27. 0:54And if I talk about on the side of
  28. 0:56governance, then I have implemented
  29. 0:58secure multi-cloud architecture,
  30. 1:01centralized access management solution
  31. 1:03aligned with compliance standard like
  32. 1:04GDPR and HIPAA. I also been involved
  33. 1:07into the client discussion and pre-sales
  34. 1:09activity where I need to translate the
  35. 1:11business requirement into the technical
  36. 1:13solutions.
  37. 1:14Along with the hands-on engineering, I
  38. 1:16have certification like AWS certified
  39. 1:18solution architect, Red Hat certified
  40. 1:20system administration, and Red Hat
  41. 1:22Ansible automation exam. So, I think
  42. 1:24yeah, this is just a brief about me.
  43. 1:27>> Okay.
  44. 1:28So, currently in your organization, are
  45. 1:30you working independently or as a part
  46. 1:32of DevOps team?
  47. 1:34>> Um like
  48. 1:36we are a service-based company, so right
  49. 1:38now I'm working on two projects
  50. 1:39simultaneously. In one project, I'm
  51. 1:41working as an individual contributor
  52. 1:43where I independently manage the DevOps
  53. 1:45activity. And in another project, I am
  54. 1:47leading a team of three DevOps engineer
  55. 1:49where I'm responsible for task planning
  56. 1:51and technical guidance.
  57. 1:53>> Okay. Okay.
  58. 1:55So, let's start with the questions.
  59. 1:58Suppose you need to host a new
  60. 2:00application that is expected to handle
  61. 2:03around 5 million concurrent users.
  62. 2:06What How would you design that? And what
  63. 2:09all services and tools would you use?
  64. 2:12>> Okay, so 5 million concurrent users. So,
  65. 2:14before designing the architecture, I
  66. 2:15would first understand the application
  67. 2:17requirement, whether it will be
  68. 2:19stateless, stateful. We need to also
  69. 2:22understand the expected traffic pattern,
  70. 2:24the database requirement, compliance,
  71. 2:25and latency expectation. And based on
  72. 2:28that, I would design a highly available
  73. 2:29and secure platform. So, like if I talk
  74. 2:31about the upper layer architecture, then
  75. 2:33it would be a it would have a route 53
  76. 2:36for DNS DNS, then we we would have
  77. 2:39CloudFront for CDN and caching. We could
  78. 2:41use a web and shield for security.
  79. 2:44Then to distribute the traffic, we can
  80. 2:45use load balancer. And if you're going
  81. 2:47with the containerization approach, then
  82. 2:49either we can go ahead with ECS or EKS.
  83. 2:51I would choose ECS if it like if the
  84. 2:55architecture is not that much complex,
  85. 2:57and ECS has lower operational overhead.
  86. 2:59But if you want some sort of advanced
  87. 3:01feature of Kubernetes like
  88. 3:03GitHubs, some sort of custom networking,
  89. 3:06or advanced deployment strategy, then in
  90. 3:08that case, I would use EKS one. And also
  91. 3:10at the same time, we can use Argo CD for
  92. 3:13the deployment.
  93. 3:14And for the backend, we can go ahead
  94. 3:16with Aurora and DynamoDB for the
  95. 3:19database, Redis for the caching, S3 for
  96. 3:21the static content, SQS Kafka for
  97. 3:24asynchronous processing. And there are
  98. 3:26multiple tools for the monitoring and
  99. 3:27CCD that we can set it up over there.
  100. 3:29So, it usually depends on the budget.
  101. 3:31Like if the client can't like can't pay
  102. 3:33it, then they can't have it. So, we need
  103. 3:34to understand the cost factor as well
  104. 3:36there over there. So, but we can go
  105. 3:38ahead with some sort of open source
  106. 3:40tools like Prometheus, Grafana,
  107. 3:42CloudWatch. We can go ahead with some
  108. 3:43sort of CCD tools like Jenkins, GitHub
  109. 3:46Action, GitLab CI. Yeah, so I think this
  110. 3:48would be some of my checklist that I
  111. 3:50would follow.
  112. 3:52>> Okay. Cool. Suppose you successfully
  113. 3:56build this application and your
  114. 3:58application is already in the running
  115. 3:59state, right? And you what you notice is
  116. 4:02sudden traffic spikes up to 10x, right?
  117. 4:06How would you determine whether it's the
  118. 4:08genuine traffic or a DDoS attack?
  119. 4:11>> Okay. So,
  120. 4:13as a debugging point of view, firstly I
  121. 4:15would analyze the access log to look at
  122. 4:17the source IP, request pattern,
  123. 4:19geolocation, and the request path. If
  124. 4:21the traffic is coming from diverse real
  125. 4:23users with normal behavior and valid
  126. 4:24application request, it's likely a
  127. 4:26genuine traffic spike. So, over there I
  128. 4:28would simply scale up the infrastructure
  129. 4:30using the auto scaling. But, if I see a
  130. 4:32large number of repeated requests from a
  131. 4:34suspicious IP or maybe some unusual user
  132. 4:37agent, and also at the same time if they
  133. 4:39are requesting at or targeting a single
  134. 4:41endpoint, so I would suspect a DDoS
  135. 4:43attack over there. So, to mitigate it, I
  136. 4:45would use a AWS WAF, AWS Shield, or like
  137. 4:48we can use some sort of rate limiting
  138. 4:49functions over there, and it will block
  139. 4:51the malicious IP. Like, also at the same
  140. 4:53time, we need to continuously monitor
  141. 4:55the application to ensure the legitimate
  142. 4:58users are not getting impacted over
  143. 4:59there.
  144. 5:01>> Makes sense, yeah.
  145. 5:03Um okay. So, suppose as you mentioned
  146. 5:07auto scaling in your answer, that
  147. 5:09reminds me of Kubernetes. So, have you
  148. 5:12worked on Kubernetes, yes?
  149. 5:13>> Yeah, yeah. Actually, we worked on the
  150. 5:15Kubernetes side.
  151. 5:16>> Okay. That's cool. So, suppose you have
  152. 5:19three name spaces in a Kubernetes in the
  153. 5:21same Kubernetes cluster,
  154. 5:23how would you prevent the dev name space
  155. 5:25from communicating with the prod name
  156. 5:27space?
  157. 5:28>> Okay. So, like in our case, ideally we
  158. 5:31create a different cluster for like
  159. 5:33different cluster for different
  160. 5:35environment.
  161. 5:36But, let's say if we are having all the
  162. 5:38environment into a single cluster, then
  163. 5:40over there we could use the Kubernetes
  164. 5:42network policy to restrict the cross
  165. 5:44name space communication. By default,
  166. 5:46like I would deny all the ingress and
  167. 5:48egress traffic, then explicitly we can
  168. 5:50allow only the required communication
  169. 5:52within the namespace that is required
  170. 5:54over there.
  171. 5:56>> Got it. Got it.
  172. 5:57So, have you created any, you know,
  173. 6:00deployments or services or ingress out
  174. 6:02there from the YAML file? Have you ever
  175. 6:05>> like we are creating it from the scratch
  176. 6:06even though we are working into the helm
  177. 6:08chart format, so yes.
  178. 6:10>> Okay. So, just let me know what happens
  179. 6:13internally when you run a kubectl apply
  180. 6:16hyphen f the supposed deployment.yaml.
  181. 6:18What happens internally?
  182. 6:20>> So, whenever we apply the kubectl
  183. 6:21command, so basically kubectl reads the
  184. 6:23YAML for manifest, then it sends the
  185. 6:25request to the kube API server, and then
  186. 6:28API server validates the request, some
  187. 6:30sort of authentication and
  188. 6:31authorization, and then it store the
  189. 6:33data into the etcd.
  190. 6:36Then the scheduler assign the port to a
  191. 6:38suitable worker node, then the kubelet
  192. 6:40on that specific node receives the
  193. 6:41instructions and ask the container
  194. 6:43runtime to create the container.
  195. 6:45Finally, the port start running and the
  196. 6:46container, like, sorry, the controller
  197. 6:48continuously ensure that the desired
  198. 6:50state is maintained, like, whatever the
  199. 6:51replica that you are mentioning inside
  200. 6:52the deployment file.
  201. 6:54So, usually this is how the flow works.
  202. 6:57>> Okay. So, based on your previous answer,
  203. 7:01um I assume that you're pretty much well
  204. 7:05versed with the Kubernetes architecture,
  205. 7:07no? Like, you have Have you heard about
  206. 7:10etcd on the master plane?
  207. 7:12>> Yes, yes. So, basically it's a database
  208. 7:13where it store the data of everything.
  209. 7:16>> Okay. So, like,
  210. 7:19answer this, like, why did Kubernetes
  211. 7:22choose etcd instead of any other
  212. 7:24database? They could have chosen any
  213. 7:25database, but why etcd?
  214. 7:29>> Um
  215. 7:30like, to be honest, like, I haven't
  216. 7:32looked into that depth, but my
  217. 7:34understanding is that etcd is built for
  218. 7:35distributed system, and it provides a
  219. 7:38consistency and leader election,
  220. 7:39specially the leader election.
  221. 7:41So, which are important for Kubernetes,
  222. 7:42so that is what my current
  223. 7:44understanding, but I would
  224. 7:45uh need to check into a bit more. Yeah.
  225. 7:49>> So, yeah, fine, fine, fine. Makes sense.
  226. 7:51So,
  227. 7:52>> [clears throat]
  228. 7:53>> your your you must have hosted, you
  229. 7:55know, microservice uh applications or
  230. 7:57the websites out there, right?
  231. 7:59>> Mhm.
  232. 8:00>> So, what do you think like how does the
  233. 8:03request reach your pod when a user
  234. 8:04accesses your website?
  235. 8:07>> So, like whenever user sends a request
  236. 8:08to our website, the request first
  237. 8:10reaches our reaches to the DNS. So, it
  238. 8:13can be route 53, it can be GoDaddy, it
  239. 8:15can be anything, and then it resolves
  240. 8:17the domain from there. And then it
  241. 8:18reaches like if you are using the
  242. 8:20application load balancer, then it will
  243. 8:21firstly go to the application load
  244. 8:23balancer in our case.
  245. 8:24Then it forwards the request to the
  246. 8:25Kubernetes ingress controller. And the
  247. 8:27ingress controller the ingress routes
  248. 8:29the request to the appropriate
  249. 8:31Kubernetes service, and then the service
  250. 8:33forward it it to the one of the healthy
  251. 8:35pod using the kube proxy.
  252. 8:38>> Okay. Okay.
  253. 8:40Uh let me just allow me a sec. Let me
  254. 8:42have a look at your resume.
  255. 8:45Okay. So, in your achievement section,
  256. 8:49you mentioned that you migrated 500 TBs
  257. 8:52of S3 bucket of data of S3 bucket from
  258. 8:55one AWS account to another.
  259. 8:58Mhm. Um like from the point of
  260. 9:01accomplishments, isn't that
  261. 9:02straightforward? You you could have just
  262. 9:05used AWS data sync, right?
  263. 9:08Can you just elaborate on that?
  264. 9:10>> So, even though to perform that
  265. 9:11migration activity, we used the AWS data
  266. 9:13sync. So, the migration part was never a
  267. 9:16bigger challenge for us. The real
  268. 9:18challenge was to how to very re-verify
  269. 9:21that all the data has been migrated
  270. 9:22successfully without missing any object
  271. 9:24because since the data was into the 500
  272. 9:26TB of data, so it was very difficult for
  273. 9:28us to re-verify the objects whether
  274. 9:30everything has been successfully
  275. 9:31migrated or not. So, to solve that
  276. 9:34specific issue, I I the some more
  277. 9:36services of AWS like I used the S3
  278. 9:38inventory. It basically generates a CSV
  279. 9:40format containing all the objects in a
  280. 9:42in the bucket. And I then loaded that
  281. 9:44inventory into a DynamoDB table and like
  282. 9:47then simply compared the source and the
  283. 9:49destination inventory to identify any
  284. 9:51missing object with the help of that
  285. 9:52specific DynamoDB database.
  286. 9:55>> Okay. Okay. Makes sense.
  287. 9:58Um okay. So, how like uh in your 6-7
  288. 10:02years of DevOps career, I hope that you
  289. 10:06must have built a lot of pipelines out
  290. 10:08there, right?
  291. 10:09>> Correct.
  292. 10:10>> Uh so, can you just walk me through the
  293. 10:12best CI/CD pipeline that you built?
  294. 10:16>> So, actually we I worked on multiple
  295. 10:17pipelines with different types of tools
  296. 10:19and different different type of
  297. 10:20technology, but in my recent project
  298. 10:22that currently where I'm working on, so
  299. 10:24let me walk you through over there. So,
  300. 10:26basically the approach would be almost
  301. 10:28similar for almost all the pipeline all
  302. 10:29the tools. So, usually whenever a
  303. 10:31developer pushes the code, the pipeline
  304. 10:33will automatically get triggered. It
  305. 10:35will firstly pull the latest code from
  306. 10:36the Git. It will run some unit test
  307. 10:38cases. Then it will perform some sort of
  308. 10:40code quality analysis using the
  309. 10:41SonarQube. Then while creating the
  310. 10:43image, we also use some sort of
  311. 10:45vulnerability scanning tool like Trivy
  312. 10:46and everything. Then it pushes the image
  313. 10:48to the ECR.
  314. 10:50Then for the CD part, we are having
  315. 10:52different repositories. So, over there
  316. 10:53we update the image tag in the GitHub
  317. 10:55repository.
  318. 10:57Then Argo CD automatically detect the
  319. 10:59change and it synchronize it with the
  320. 11:01our Kubernetes cluster.
  321. 11:03And then the application was uh in our
  322. 11:05in our current project, the application
  323. 11:06was deployed using the rolling update
  324. 11:08strategy. And post deployment, health
  325. 11:10checks were performed before making
  326. 11:11before marking the deployment as
  327. 11:13successful.
  328. 11:15>> Uh you mentioned uh the application was
  329. 11:17deployed using rolling update, right?
  330. 11:20>> Correct.
  331. 11:21>> Uh so, why didn't you use a canary
  332. 11:23deployment and instead you used rolling
  333. 11:25update? Because canary deployment is too
  334. 11:27much in fashion, no?
  335. 11:28>> Correct.
  336. 11:29>> So, why?
  337. 11:30Yeah.
  338. 11:30>> So, basically the thing is that we are
  339. 11:32at the stage where we are enhancing the
  340. 11:34infrastructure.
  341. 11:35So, like in the current project, like we
  342. 11:37are using the rolling update because
  343. 11:39like the application was backward
  344. 11:40compatibility, and the deployment risk
  345. 11:42was very low. So, it provided us a
  346. 11:44zero-downtime deployment with minimal
  347. 11:45operational overhead.
  348. 11:47But let's say if we have to set up a
  349. 11:48canary deployment for some sort of
  350. 11:51business-critical application or major
  351. 11:52releases, we could really
  352. 11:54implement the canary deployment. It just
  353. 11:56like we need to just
  354. 11:58do some sort of basic testing. Like the
  355. 12:00basic implementation, it will just allow
  356. 12:01us to allow give us a feature to
  357. 12:03release a new version to a small
  358. 12:05percentage of user. Then we can simply
  359. 12:07monitor the key matrices like the error
  360. 12:08rate, latency, and then we can simply
  361. 12:10gradual gradually increase the traffic.
  362. 12:12And if any issue is detected over there,
  363. 12:14then we can simply roll back into the
  364. 12:16previous version. So, if if needed, then
  365. 12:18we can also implement the canary
  366. 12:19deployment. But we are at the stage
  367. 12:21where we are enhancing the
  368. 12:22infrastructure.
  369. 12:24>> Got it. Got it. Makes sense. Yeah.
  370. 12:27So, like while creating CI/CD
  371. 12:30deployments out there, there you might
  372. 12:32have faced a lot of challenges.
  373. 12:34Sometimes some havoc, some
  374. 12:37issue in the configuration. So, let me
  375. 12:39give you a scenario. Let's say a
  376. 12:41developer accidentally commits an AWS
  377. 12:44access key
  378. 12:45and the secret key to the Git Git
  379. 12:47repository.
  380. 12:49What And that that's a blunder, no? What
  381. 12:52immediate actions would you take and how
  382. 12:55would you prevent this from happening in
  383. 12:57future?
  384. 12:58>> So, firstly, what I will do, I will just
  385. 13:00simply disable or rotate the
  386. 13:02compromised access key so that it cannot
  387. 13:04be used anymore. And then I will go to
  388. 13:07the AWS, and then we can we could check
  389. 13:09the cloud trail to just to identify if
  390. 13:11any suspicious activity has been done
  391. 13:13with the those compromised key or not.
  392. 13:16And then
  393. 13:17for the future, we can integrate some
  394. 13:19sort of Git secret or Git leak into the
  395. 13:21CI pipeline to detect whether any of the
  396. 13:23secret is getting pushed or not before
  397. 13:25the code is getting merged.
  398. 13:26And then like at the like the best
  399. 13:28practices is we should store some sort
  400. 13:30of sensitive credential into the AWS
  401. 13:33secret manager or we could save it to
  402. 13:34HashiCorp Vault as well.
  403. 13:36And also, at the same time, we need to
  404. 13:38educate the developers on secure
  405. 13:40credential management as a part of the
  406. 13:42development workflow.
  407. 13:44>> Okay. Okay, makes sense.
  408. 13:47So, sup- let's suppose
  409. 13:51like by using your CI/CD pipeline or
  410. 13:54something, uh you deployed a new version
  411. 13:56of the application, right? The
  412. 13:58deployment is showing successful, that
  413. 14:00means the green deployment, right? So,
  414. 14:02it's green on your side, but the
  415. 14:04customers are, you know, getting a lot
  416. 14:06of 502 bad gateway errors.
  417. 14:09Right? So, how would you troubleshoot
  418. 14:10that?
  419. 14:12>> So, firstly,
  420. 14:14uh okay, the deployment goes green, but
  421. 14:15still the users are impacting or getting
  422. 14:17502 errors.
  423. 14:18>> Yeah, 502. Yeah, yeah.
  424. 14:19>> So, maybe firstly, I will check the logs
  425. 14:21of the pod, and maybe I will check
  426. 14:23whether that pods are up and running or
  427. 14:25not. Maybe I would also check some sort
  428. 14:26of readiness and liveness probe as well.
  429. 14:29And then, I will verify whether the
  430. 14:30service is correctly pointing or
  431. 14:31forwarding the traffic to the healthy
  432. 14:33pods or not.
  433. 14:34Then, even though if that is also
  434. 14:35correct, then I will check the ingress
  435. 14:37controller or the application at the LB
  436. 14:39logs to identify whether the from where
  437. 14:41the 502 is being generated. Then, also I
  438. 14:44will at the same time, I will I will try
  439. 14:45to check some sort of application logs
  440. 14:47for any sort of startup or runtime
  441. 14:49error.
  442. 14:50Then, like, I will also ensure that the
  443. 14:52the application is listening on the
  444. 14:53correct port and the service target port
  445. 14:56matches the container port that we
  446. 14:58defined inside the deployment file.
  447. 15:00And even though if the issue is still
  448. 15:02persistent over there, so my first
  449. 15:04preference would be to roll back it to
  450. 15:05the
  451. 15:06previous stable version to restore the
  452. 15:08services while investigating the root
  453. 15:10cause. And we can we could do it onto
  454. 15:12the lower environment. So, firstly, I
  455. 15:13will roll roll back it to the previous
  456. 15:15version, so that at least the end
  457. 15:16customer would not impact anymore.
  458. 15:19>> Got it. So, like, suppose as you
  459. 15:22answered, you will check this you will
  460. 15:24check a lot of things. So, like during
  461. 15:27your investigation of the checks, you
  462. 15:30find that the suppose you during your
  463. 15:32investigation of the checks, you find
  464. 15:33that the 502 error is caused by the
  465. 15:36database, right? And the database CPU is
  466. 15:39consistently above 90% since you made
  467. 15:44the deployment. And one there's one
  468. 15:46query that is being executed hundreds of
  469. 15:48times per minute because of the because
  470. 15:52of the particular application
  471. 15:53functionality that's assume. So, how
  472. 15:55would you solve this?
  473. 15:57>> So, the firstly I would do is to confirm
  474. 16:00whether the repeated query is actually
  475. 16:01necessary or it's an application issue.
  476. 16:04Let's say if the data doesn't change
  477. 16:05frequently like whenever user is
  478. 16:07requesting that specific query. So, I
  479. 16:09would introduce Redis as a caching
  480. 16:11layer. So, instead of hitting the
  481. 16:12database directly for every request, the
  482. 16:14application firstly it would go to the
  483. 16:16Redis and if the data is available it
  484. 16:17would return the response directly from
  485. 16:20there. If the data is not present in
  486. 16:21Redis, it would query the database and
  487. 16:23then store the result in Redis for the
  488. 16:25subsequent request. So, like this would
  489. 16:27significantly reduce the number of
  490. 16:30database queries. It will also reduce
  491. 16:32the CPU utilization. So, maybe this is
  492. 16:35how we could improve the response time
  493. 16:37for the request.
  494. 16:39>> Okay.
  495. 16:41Okay.
  496. 16:42So, have you worked on Terraform?
  497. 16:44>> Yes, I'm working on Terraform. Yeah.
  498. 16:46>> Okay. So, how would you structure
  499. 16:49Terraform code for large enterprise
  500. 16:50project with multiple environments like
  501. 16:53dev, prod, and stage?
  502. 16:56>> So, we can create a modular structure in
  503. 16:58which like we can create separate module
  504. 17:00folder for component like VPC, EKS, RDS,
  505. 17:03or whatever the services that we are
  506. 17:05using into our current project. And then
  507. 17:07there would be a separate environment
  508. 17:08folder such as dev, stage, prod in which
  509. 17:10where each environment will call the
  510. 17:12same module with the different variable
  511. 17:14values.
  512. 17:15So, it would like also at the same time
  513. 17:16each environment would have its own back
  514. 17:18end and state file to keep the interest
  515. 17:20isolated. And then we could simply
  516. 17:22integrated that Terraform with a CCD
  517. 17:24pipeline to so that every infrastructure
  518. 17:26changes goes through review, validation,
  519. 17:28and approval being before being getting
  520. 17:30applied over there.
  521. 17:31So this is what we can do here.
  522. 17:33>> Okay. So while working with Terraform,
  523. 17:36Terraform being too sensitive and you
  524. 17:39know, you you need to do your work in a
  525. 17:41good fashion, right? So
  526. 17:45you must have faced some kind of you
  527. 17:47know, issue while issue as in like
  528. 17:50suppose I'll give you an example. Like
  529. 17:52what like what would happen if someone
  530. 17:54manually changes the infrastructure in
  531. 17:56AWS outside of Terraform? What would
  532. 17:59what like what happens in that case?
  533. 18:01>> Let's say
  534. 18:02if they made some sort of changes into
  535. 18:03the AWS cloud and if we are not having
  536. 18:05that code into our Terraform, then
  537. 18:07whenever we will try to run the
  538. 18:08Terraform plan, the Terraform will
  539. 18:10detect the drift and shows the
  540. 18:11difference between the desired state and
  541. 18:13the current state. So like it depend on
  542. 18:15the change. Like if the manual change is
  543. 18:16not part of the Terraform code,
  544. 18:17Terraform apply may revert it back to
  545. 18:19the desired state. But let's say if the
  546. 18:21manual change is needed into our current
  547. 18:23code right now. So I would firstly
  548. 18:25review it and then update the Terraform
  549. 18:26code so that the infrastructure and the
  550. 18:28code remains in sync. So this is what
  551. 18:30this is how we can solve it.
  552. 18:33>> Okay.
  553. 18:34So
  554. 18:35when do you use depends on? Can you just
  555. 18:37give me an example?
  556. 18:39>> So like
  557. 18:40I use depends on when I want to make
  558. 18:43one resource is created only after
  559. 18:46another resource is ready. So there is a
  560. 18:47dependency. Over there we could get So
  561. 18:49if I give you a small example, then
  562. 18:51let's say we create a private route
  563. 18:53while creating VPC, right? So I want to
  564. 18:56ensure that the net gateway because
  565. 18:58inside the route we attach the net
  566. 18:59gateway to the routes. So I want to
  567. 19:01ensure that the net gateway is fully
  568. 19:02created before the route points to it.
  569. 19:04So in such case I can accidentally use
  570. 19:06the depends on to guarantee that the
  571. 19:08creation order.
  572. 19:10>> Okay. Okay.
  573. 19:12So let's come back to AWS maybe just a
  574. 19:15few questions on AWS.
  575. 19:17What would happen if an entire AWS
  576. 19:21region goes down and how would you
  577. 19:23recover your application?
  578. 19:25>> Like we are already deploying our
  579. 19:27application with the help of
  580. 19:28multi-region. But even though in our
  581. 19:31current project since we are working
  582. 19:32with US based client, so we perform some
  583. 19:35sort of disaster recovery drills every 3
  584. 19:37months to ensure our recovery process
  585. 19:39actually works. So, there are multiple
  586. 19:41steps that we need to follow. So, we are
  587. 19:42keeping a backup of our database like we
  588. 19:45use the database snapshot. It will
  589. 19:47automatically copy like it will
  590. 19:49automatically create the snapshot and
  591. 19:50there is lambda function which will move
  592. 19:51the snapshot from one region to another
  593. 19:53one another one, sorry. And then we also
  594. 19:56enable the S3 cross region application
  595. 19:57that is simply CRR so that our
  596. 19:59application data is already available in
  597. 20:01the DR region. And since our
  598. 20:03infrastructure is managed through
  599. 20:04Terraform, so we can quickly recreate
  600. 20:06the infrastructure in the secondary
  601. 20:07region. That won't be an issue much
  602. 20:08more.
  603. 20:09And once the application is up and
  604. 20:11ready, like we can simply update the
  605. 20:12route 53 to route the traffic to the
  606. 20:14healthy region.
  607. 20:15And we have already tested this DR
  608. 20:17process multiple times in our current
  609. 20:19project. So, like right now like we are
  610. 20:22much confident that we could cover the
  611. 20:24application like we can recover the
  612. 20:25application within the agreed RTO or the
  613. 20:28RPO instead of figuring it out at the
  614. 20:30run time.
  615. 20:31>> Okay.
  616. 20:33Okay, okay.
  617. 20:34I was just going through your resume. I
  618. 20:37find one thing very interesting.
  619. 20:40You worked on
  620. 20:43>> [clears throat]
  621. 20:43>> HIPAA and GDPR compliances, right?
  622. 20:46>> Correct.
  623. 20:47>> So, what was your exact role in this?
  624. 20:50Can you just mention?
  625. 20:51>> So, basically I was working for a client
  626. 20:53and they are into the health care
  627. 20:54sector. So, since they are applying
  628. 20:56around the data of the patient. So, they
  629. 21:00need to be a bit more
  630. 21:02concerned about the security part. So,
  631. 21:04like currently they They into the Europe
  632. 21:06under the GDPR compliances. So, when
  633. 21:07they expanded their business to the US
  634. 21:09so they also needed to comply with the
  635. 21:11HIPAA requirement. So, as a part of the
  636. 21:13initiative, our our focus was on
  637. 21:15strengthening the infrastructure from a
  638. 21:17security and compliance perspective. So,
  639. 21:20if I give you a small example, there
  640. 21:21were a lot a lot of checklist over
  641. 21:23there, but if I give you a small
  642. 21:24example, then uh we ensured the private
  643. 21:26S3 bucket were accessed through VPC
  644. 21:28endpoint instead of the public internet.
  645. 21:30Then we are implementing some sort of
  646. 21:31governance policy in our community
  647. 21:32cluster to uh to enforce the secure
  648. 21:35deployment uh standards. Then we make
  649. 21:37sure that the PI and the PHI data was
  650. 21:39never written to the application logs by
  651. 21:42masking or removing the sensitive
  652. 21:43information from that logs.
  653. 21:45Then we are also like we also enforce
  654. 21:47the encryption. We also use some sort of
  655. 21:49least privileges I am access and
  656. 21:51regularly we review the security and
  657. 21:53everything, yeah.
  658. 21:55>> Okay. Okay.
  659. 21:57So, let [clears throat] me give you a
  660. 21:59real-life scenario, right?
  661. 22:01>> Mhm.
  662. 22:01>> So, uh right now FIFA is uh too much
  663. 22:05trending and the ads and the banners in
  664. 22:07the FIFA. You must have heard that the
  665. 22:09in different countries they show
  666. 22:11different kind of ads or different kind
  667. 22:12of pages, right? So,
  668. 22:15>> [clears throat]
  669. 22:15>> like I would like to question on that in
  670. 22:17CloudFront, if you want to show
  671. 22:19different pages based on the user's
  672. 22:21country, how would you implement that?
  673. 22:25>> Um different pages based on the their
  674. 22:28country, right?
  675. 22:29>> Yeah.
  676. 22:30In CloudFront.
  677. 22:31>> Oh, CloudFront, right.
  678. 22:32So, there would be multiple ways of
  679. 22:34doing it. Uh like the firstly, the
  680. 22:36simplest would be like we can write up a
  681. 22:37small logic inside the application that
  682. 22:40will handle the uh
  683. 22:42logic which will uh move the traffic
  684. 22:44from one page to different one according
  685. 22:46to the user's country. But let's if you
  686. 22:48are specifically going with the
  687. 22:50CloudFront, so there is a there is one
  688. 22:51more option that we can do it uh
  689. 22:53CloudFront
  690. 22:55provides us the CloudFront function.
  691. 22:57So, we could use the CloudFront
  692. 22:58functions over there and there is a uh
  693. 23:00header like we can use uh, viewer
  694. 23:02country header, like I don't exactly
  695. 23:04remember the header, but there is some
  696. 23:05sort of CloudFront viewer country, some
  697. 23:07sort of header like that.
  698. 23:07>> Yeah.
  699. 23:08>> And then we can write up a small
  700. 23:09CloudFront function, and that it will
  701. 23:10automatically inspect that specific
  702. 23:12header, and rewrite the request URI or
  703. 23:14the route the users to a different uh,
  704. 23:17content over there. Let's say if the
  705. 23:18user is coming from India, so it will
  706. 23:20uh,
  707. 23:21access on the server page like with some
  708. 23:23sort of landing page. So we could create
  709. 23:25a CloudFront function. Then there are
  710. 23:27multiple ways of doing it. So even
  711. 23:28though we could implement some sort of
  712. 23:29Lambda function through which
  713. 23:30Lambda@Edge which will handle this
  714. 23:32specific logic. So there are multiple
  715. 23:34ways of doing it.
  716. 23:36>> Mhm. Okay.
  717. 23:38Okay.
  718. 23:40So
  719. 23:41uh, have you worked in multi-cloud
  720. 23:42environments?
  721. 23:44>> Yeah, so like uh, like within this uh, 7
  722. 23:47or 6 years of my cloud journey, so I got
  723. 23:50a chance to work on different different
  724. 23:51projects. I worked on AWS, GCP, Azure. I
  725. 23:54did a migration of like multiple times I
  726. 23:56did a migration of from GCP to AWS as
  727. 24:00well. I did a lot of uh, I have created
  728. 24:03a lot of pipeline in terms of Azure
  729. 24:04DevOps as well. So like it's it's like
  730. 24:06they have worked on almost all the cloud
  731. 24:08provider, but most extensively on the
  732. 24:10side of AWS.
  733. 24:12>> Most ex- And I also see that you have
  734. 24:16uh,
  735. 24:17AWS certification, right?
  736. 24:18>> Correct. Correct. So yeah, it's a
  737. 24:19professional one.
  738. 24:21>> Yeah, it's a professional one. So when
  739. 24:23did you get this?
  740. 24:25>> So I did that in the year somewhere
  741. 24:27around 2023. So maybe it would be in
  742. 24:30some sort of expiration state or maybe
  743. 24:32like it would be a bit closer to the
  744. 24:34expiration.
  745. 24:35>> Yeah, it's like fine. It's like
  746. 24:38Okay.
  747. 24:39Uh, let me go through your resume once
  748. 24:42more.
  749. 24:43Just allow me a second.
  750. 24:45Okay. So I was Yeah, I was asking that
  751. 24:49Mhm. Like, [clears throat] can you just
  752. 24:51tell me more about Kibana, like a little
  753. 24:53about Kibana?
  754. 24:55>> So right now we have used the Kibana in
  755. 24:56our current project, so it is helping us
  756. 24:58to uh,
  757. 25:00strengthening the security of the cabin.
  758. 25:01So, let me give you a small example. So,
  759. 25:03we have like it's just a YML file in
  760. 25:04which you could create any sort of
  761. 25:06policy that you wanted to use in your
  762. 25:07Kubernetes cluster. In our case, like
  763. 25:09I'm using like this is one of the policy
  764. 25:11that we are using.
  765. 25:12So, that we have mentioned the account
  766. 25:14ID of the AWS account account AWS
  767. 25:15account ID. So, that whatever the image
  768. 25:18that is present into that specific AWS
  769. 25:20account ID, only those deployment would
  770. 25:22be able to deploy. They won't they won't
  771. 25:24be able to deploy any sort of basic like
  772. 25:26Ubuntu, Nginx, Python. They won't be
  773. 25:28able to deploy anything. Only the images
  774. 25:30that is present on that specific AWS
  775. 25:32account ID. So, this is how we are
  776. 25:34restricting to use
  777. 25:36secure images that is present on our AWS
  778. 25:38account. Yeah.
  779. 25:40>> Okay.
  780. 25:42Okay, makes sense. Um
  781. 25:45yes, I guess I covered most of the
  782. 25:47points.
  783. 25:49Um yes, so yep, I'm good from my side.
  784. 25:53Um maybe if you want to have some
  785. 25:54questions, please spam them.
  786. 25:56>> So, actually I don't have any specific
  787. 25:58question, but just wanted to understand
  788. 25:59that what kind of project or business
  789. 26:01domain that your organization is working
  790. 26:02on and also maybe some bit idea about
  791. 26:04the tools and the technology.
  792. 26:07>> Okay. So, in our organization, we work
  793. 26:11with multiple clients across multiple
  794. 26:12domains such as, you know, health care,
  795. 26:15fintech, a lot of e-commerce clients out
  796. 26:17there, retail clients out there. So, the
  797. 26:20technology stack varies like depending
  798. 26:23on the client based on the client
  799. 26:24requirement. But most of our project
  800. 26:26revolve around AWS, Kubernetes,
  801. 26:28Terraform, Docker, GCP, and monitoring
  802. 26:30tools.
  803. 26:31So, as a DevOps engineer, you'll likely
  804. 26:34work on different projects and, you
  805. 26:37know, get exposure to multiple
  806. 26:38technologies over the same time. Yeah.
  807. 26:41I guess I answered your question.
  808. 26:43>> Yeah, yeah. Yep, perfect. I think so,
  809. 26:45yeah, that is the only question that I
  810. 26:47have right now.
  811. 26:48>> Okay. Okay.
  812. 26:50So, nice meeting you, Jadeja. I'll just
  813. 26:53pass on my feedback to the HR. She'll
  814. 26:55get back to you.
  815. 26:55>> Sure. Thank you so much. Thank you so
  816. 26:57much. Yeah, have a nice day.
  817. 26:58>> Thank you.

About this transcript

This page contains the full transcript of DevOps Live Interview | DevOps | Cloud | DevSecOps | SRE #interview #devops #aws #docker #cloud by jadeja rajpal sinh, generated from the public captions YouTube serves with the video. The transcript has 4,957 words across 817 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.