Lecture 03: Sorting, Searching and Arrays — Transcript
Full transcript
- 0:14[Music]
- 0:15[Music]
- 0:16Okay. So meanwhile we can just reiterate uh so last time we discussed
- 0:25uh time complexity of insertion sort and merge sort and then we started uh some examples of DS.
- 0:44Yes. So in DS examples we did uh uh basically binary search
- 0:54both sequential and fast and uh we introduced this problem of range minima for which we haven't
- 1:04given a good solution. So we'll do this in the future. So let's now continue with uh
- 1:13a problem that uh we have seen before and some of
- 1:18you asked how we will solve it so long we'll give a fast algorithm.
- 1:26So we saw this iterative algorithm iterative fib
- 1:34which takes uh n steps and same space
- 1:50right so using the using an array which will store the values of previous fi minus on you
- 1:56can calculate f_subi and that you can do. So you can calculate the nth fibonaki number in
- 2:05around n steps which is much much faster than as we have seen recursive algorithm.
- 2:13But uh I gave you this question of uh calculating
- 2:25in a super fast way
- 2:31for fn mod 2024.
- 2:41Okay. Okay, we can call this number m. I mean this is the current year but next year it will be
- 2:47something else. The point is that this uh whatever is the year you want to calculate fn mod m
- 2:55uh so m will be very small right this is just a fourdigit integer uh but f of n is extremely
- 3:03large f of n has more than n I mean it is it is more than 1.5 to the n so it has more than
- 3:09uh it has around n digits and assume n to be very large as we said n can be long int or long int
- 3:18So it will be 64 bits itself. So you this iterative fib will not work for this. The program
- 3:25will never stop. End steps will be too much. Right? So any ideas how you will calculate this
- 3:35is yeah. So matrix multiplication is what somebody says. Uh it'll be even better than that. We will
- 3:42actually convert this algorithm into a matrix powering problem. Okay. Okay. So you have to
- 3:48ultimately I will design a matrix whose uh nth power mod m will give you the answer
- 3:57and somehow that nth power of a number or a matrix there are you can give a super fast
- 4:03algorithm. Okay. So super fast means login. So it will work for n which is long login also.
- 4:17Yeah, so that's what we want to do in the end.
- 4:22Uh but yeah, let's just uh for completeness or for revision look at what uh the older
- 4:28algorithms will do. So this iterative fib now it takes two uh arguments,
- 4:39right? N is the number which is extremely large and M is the M is for example 2024.
- 4:45It's very small. You just want to calculate the remainder. So how will you modify 8 fib
- 4:51n? So the addition will now be mod m, right? You'll just store the remainder.
- 5:04So same uh pseudo code and same array.
- 5:18But now the arithmetic the uh addition is mod m.
- 5:28Okay, which is uh an array which stores even smaller integers very small. It's
- 5:33only numbers from 0 to n minus 1 or 0 to 2023. So that's the that's the
- 5:39array manipulation. In n steps you'll get you'll get the value of fn mod m.
- 5:46So let's record that. So this uh 8 fib n comma n comma n m requires
- 6:01greater than uh 3n
- 6:05uh instructions.
- 6:12I mean the instruction is just addition but then you also have
- 6:15to do uh division and compute the remainder. So it's a slightly more
- 6:20complicated instruction but since m is small this also will be fast
- 6:28and uh only so how much space is required for this how much
- 6:36are you storing? So the array size is n and uh each uh uh position in
- 6:47the array has to support uh this value up to m right. So that is n plus log m
- 7:04not plus actually this will be product
- 7:13which is also fine because log m is very small. Uh it's something like
- 7:1811. So space is only 11 11 * n. Okay. So this space is also uh I mean space is not
- 7:27the problem here. The problem will only be instructions because this uh this when
- 7:31you convert this into seconds it will be too big. So that is what is uh what is bad here.
- 7:42Yes.
- 7:48Yeah. This is just uh we are just splitting hairs. We can skip log m but
- 7:56uh if you want to see the dependence on 2024 you have to say login. So when 2024 becomes
- 8:03let's say 2 lakh 24 then the amount of numbers that you can put in a position of the array will
- 8:12be bigger. So how much bigger? So that is uh basically binary representation of m. How many
- 8:21uh what is the maximum power of two that divides m that is log m because ultimately everything in
- 8:29a computer is stored as 01 right I mean if you are from electrical engineering you know that
- 8:36the only thing a circuit can store is on or off state. So everything is ultimately reduced to 01.
- 8:43How many 01s will 2024 require? So that in binary representation is optimal that is
- 8:49login. But that here is very small. So we can potentially skip it. Uh
- 9:00so yeah the time is linear but it's still bad because our n is going to
- 9:04be very large. So it doesn't good give a good uh time in seconds.
- 9:12And uh yeah again just for completeness let's also check recursive fib
- 9:21which is now n comma m.
- 9:26So this was uh if n is large
- 9:34then u you recurse
- 9:42so which is uh n - m and n - 2
- 9:58and with you have to do mod m arithmetic again. So here recursive fib n minus 1
- 10:05comma m will give you a number between 0 to n minus one and recursive fib n minus2 m will
- 10:11again give a number 0 to n minus one and when you add the two this can be bigger than m. So
- 10:16you have to again take remainder. Okay. So this is again arithmetic being done. this ar this is
- 10:21the instruction. So two calls of recursion and then uh one arithmetic instruction. So
- 10:29that will give you the the answer and the boundary condition you handle as base case
- 10:39which is you just have to output n 0 or one when n is 0 or one respectively. So this is
- 10:47uh much worse as expected. So this takes uh
- 11:02um this takes at least uh as many steps as is the value of f of n the nibaki number.
- 11:17Right? Because the recurrence that you get for time will be the same. It's the time for
- 11:23n minus one plus time for n minus2 plus the uh additional arithmetic plus mod m.
- 11:34So
- 11:37which is significantly bigger than 1.5 to the n. So this is really useless.
- 11:43um or somebody corrected it to n minus one I think. So it is it is growing exponentially
- 11:56and n already was uh very large. So this is
- 11:59in the exponent n is sitting. So this is even more infinite time.
- 12:07So yeah, now let us come back back to the super fast algorithm that we want
- 12:13to do in this class. Uh so that's the third idea,
- 12:21right? So that's a clever insight which many of you have not seen probably.
- 12:33So the clever insight in many of the algorithms in this course will be uh do double instead of + one.
- 12:52Okay. So what does that mean? So instead of taking one step you take double steps. Uh so for example
- 13:03in the recurrence uh this uh f_subi equal to f_i -1 plus fi -2 or even the recursion you are
- 13:13reducing the problem of n to n minus one right or in other words the opposite way you are going from
- 13:18n minus 1 n -2 to n so you are doing a plus one so instead of doing a plus one you should uh actually
- 13:25go from n to 2 n you should double which uh in the opposite way means that if you want to calculate
- 13:33n you should use n by two. Okay. So instead of + one you do double or instead of minus one you have
- 13:42that will be the trick. um and will of of course I mean then what you have to do what you have to see
- 13:50is you have to see the recursion in a completely new way right so fibuachi sequence by definition
- 13:57is this is there a way to uh reinterpret this so that you can do the doubling thing or havinging
- 14:09thing you want to relate n to or i to i by2 Yeah. So you have to really change the language
- 14:16in which you are doing this calculation. Um so let us change the language to a matrix.
- 14:28Yeah. And why do we do that or how did we get this idea? This I don't know this you have
- 14:32to ask the student. This is very mysterious right? Why should you look at a matrix when
- 14:38you are only interested in numbers right? So that deep inside the student will tell
- 14:43you but I can give give you the solution. So the solution is that you write it like this.
- 14:59So you look at the evolution of uh I -1 I -2
- 15:04uh locations in the sequence to the next locations which is I and
- 15:10I minus one. Right? So the evolution is given by a matrix. What is the matrix?
- 15:17So this is the 1 one row and this is the one zero row.
- 15:23Okay. So this is the evolution of uh two entries in the sequence instead of one
- 15:29entry. So that's the matrix representation of the recurrence. It's completely equivalent.
- 15:39So evolution of f as a matrix
- 15:45and in this case 2x2 matrix. So if you look at this evolution then the advantage is
- 15:52u is the following. So we can call this matrix A
- 16:06and uh we can apply this again right so now from
- 16:11I - 1 I - 2 we can go to I - 2 I - 3 and what will happen
- 16:22so what what is the matrix here.
- 16:27Yeah, because the matrix notice that the matrix here is a absolute constant. It
- 16:33doesn't depend on uh I or N or whatever. So this will be this can be repeated. So
- 16:39you will get a square here. And now with little imagination you can see
- 16:46where we are headed. So you will get e to the n minus1 * f_sub_1 and f_sub_0.
- 16:56Uh this is in i notation
- 17:06is that correct?
- 17:09So this is uh then a to the i - 1 * 1 0. So what we have learned from here is that uh
- 17:24f of n is equal to
- 17:31u
- 17:35the 1 comma 1 entry of a to the n minus
- 17:45Okay. So the top left entry of this mat,
- 17:48so this is a 2x2 matrix. You are powering it. So when you do this n minus one time,
- 17:54you still will have a 2 +2 matrix with very big numbers as the four entries. So the top
- 18:00left entry is the value of f of n. Right? So same thing gives you the modern value also.
- 18:21So this is our formula. uh we take this uh trivial matrix and we exponentiate
- 18:29it n minus one times and each time we just store values four numbers
- 18:37which are from 0 to n minus one right so every time the space is very small
- 18:44uh and uh the the formula is very explicit so you can obviously do this in a for loop it is
- 18:52just multiplying two matrices is but the number of steps here is still n because the for loop has
- 19:00to compute each of these powers right so where is the advantage how do you make it super fast
- 19:09by using repeated squaring yeah so again why do we do that we don't know you have to ask the students
- 19:17so these are all deep insights right it's hard to explain why we do this but I have given you
- 19:23The main uh paradigm it is you you want to somehow double instead of just plus one. So
- 19:28instead of just doing a for loop uh with increment of plus one you want to take a
- 19:35larger leap. So what you see here is that instead of computing a square a cube from a square you can
- 19:43directly go to a to the 4. Why is that? Well if you have a square matrix you can just square it
- 19:50again. So you'll get a to the four right? So you can skip a cube and when you are at a to
- 19:55the 4 you can skip what you can skip 5 6 7 and go to eight because a4 you can just square. So
- 20:04only by doing squaring you can calculate this u I mean of course the smart ones amongst you
- 20:11will ask that what if n minus one is not a power of two what will happen then right
- 20:17so what will happen if n minus one is seven then you cannot skip seven if you skip 5 6 7
- 20:24then you then you go to eight and then you only have four and eight how do you calculate seven
- 20:32yeah So the even more smarter ones amongst you will see that what you have to do is
- 20:36n minus one you have to write in binary. So it's a sum of two powers and then only those
- 20:43two powers you have to calculate right? So seven you have to write as 1 + 2 + 4 and so
- 20:51you get a then you square it then you square it and then you multiply those. So 1 + 2 + 4
- 20:57will give you seven. So that that's how it is done. Um so let us uh note that.
- 21:10So the insight of squaring.
- 21:22Uh don't multiply when you can square.
- 21:37Right. So what you do is uh you calculate these u powers. So a a² a to the 4 a to the 8 a to the 16
- 21:53and a to the 2 to the k everything mod m right so the arithmetic always is mod m so that the
- 22:01numbers do not blow up. Uh so all the entries of this these 2x2 matrices will be very small. They
- 22:08are only between 0 to 223. Uh and there are only k steps here. Right? So for the price of k steps
- 22:18you are getting a power of 2 to k. Right? This is the power of doubling instead of plus one and
- 22:26which shows in the exponent as squaring. Right? So it's happening in the exponent.
- 22:31So this is this is actually uh the power of squaring. This is called repeated squaring.
- 22:48So yeah let us just if you have not understood the details you can work out the pseudo code.
- 23:02So let's just make some observations first. Uh so first observation is
- 23:07that we are getting to 2 to k uh power in k steps.
- 23:20Right? So this is why it is a super fast algorithm for you are reaching nth power
- 23:27in login. So it's something like binary search uh but far more complicated. It's not just matter of
- 23:36sorting and looking at the midpoint right. This is much more complicated arithmetic.
- 23:44uh so which uh means in our case that we get to
- 23:54n - 1/8
- 23:57power of the matrix equivalently we get to f of n mod m in login steps
- 24:11login or maybe just the ceiling of that the integer after login. So this is what you have
- 24:17to remember. Okay. Some more small observations we'll need. So how do you multiply matrices?
- 24:37So recall if this B if B and C are two 2 +2 matrices how many steps are
- 24:45are required? Uh no three is being too stingy. I mean the matrix itself has
- 24:56four entries. You have to see all entries. Uh it's slightly more. It's uh 4 squar + 4
- 25:074 square + 4 let me say instructions
- 25:15so assuming that this addition multiplication is for free of numbers when you have two 2 +2
- 25:23matrices you see that there are four entries each so every entry you have to multiply with
- 25:29every entry right so that's four square and uh then you have to do some additions also after you
- 25:38have multiplied you add. So it's around 4 square + 4 instructions but then each instruction itself is
- 25:48adding numbers modu and then dividing and finding the remainder model m right. So if m is growing
- 25:58then uh that is again log m another cost of log m. So it's around uh 20 log m time let me say
- 26:15so in terms of simple steps it is uh you can multiply two matrices and also calculate the
- 26:22answer mod u so the output will just be four integers between 0 to m minus one
- 26:30all that you can do in around 20 log m steps Okay. So this is considered very
- 26:35cheap. This is not our bottleneck. Uh but still you have to be careful about this.
- 26:45So how much will uh this cost a to 2 to k mod m.
- 27:01So this 2 raised to kth power we have written above right it's k steps it's
- 27:06basically k squarings each squaring you will do like b cross c subine you will call that
- 27:14subruine k times correct so it's that answer time k so let me put it like that 20k log m
- 27:30okay this is the complex lexity of calculating a to 2 tok mod. So these
- 27:36are the basics of our pseudo code right do repeated squaring each squaring costs around
- 27:46log m and then you repeat this k times so you will get to a to 2 to k mod m.
- 27:56Okay. So any any questions at this point? So whatever you
- 28:00did not understand here you check as an exercise
- 28:06because these are the things which will be
- 28:08uh which I will just assume and you'll be asked questions in the assessments
- 28:14about this. Okay. These are the very basics of uh pseudo code analysis.
- 28:24So now let us write down this uh clever fib
- 28:32n comma m.
- 28:36Uh
- 28:40maybe one one more remark can be made here. What is the space which is required for this
- 28:44calculation? A to 2 k mod m how much space is needed? So each time you only have to store
- 28:54four numbers and that two mod m right. So the space is just four times log m.
- 29:07So the space here is trivial. This is just uh m is 2024. So it's around log m is around 11. So 44
- 29:17uh bits are are all you need to store. Plus there will be some overhead of your of your
- 29:23device. Uh but removing that overhead your requirement is only 44 or 50 bits, right?
- 29:38Yeah. Every space can be reused. And first of all I mean there is no recursion here.
- 29:47It's a for loop which is just going from E4 816 to 2. Every time the same space you
- 29:53can reuse the previous space you are just uh whatever is written you are squaring. Right?
- 29:59So the space here is uh is almost zero. There is no space requirement in this algorithm. So
- 30:07it's a super fast algorithm with almost zero space. Right. So these are dream algorithms.
- 30:17Okay. So yeah, let us get into more details of the pseudo code. [Music] So
- 30:27we'll keep the we'll keep uh we'll maintain an array s which will store your matrices.
- 30:42Yeah, actually this algorithm will have more space
- 30:44requirement but we'll see that later. The previous statement was still right.
- 30:51So let's uh initialize the array with the matrix A which is defined if you remember 1 1 0
- 31:02and let us take K here to be
- 31:08K we will take log of N minus one ceiling.
- 31:17So this is the maximum power of two that we want to reach right we want
- 31:21to calculate a to n minus one. So uh s0 I have substituted to be a and uh
- 31:37so as I have said before n minus one you have to first write in binary so that you know
- 31:43which two powers are important. So n minus one will define this binary representation
- 32:00or uh won't even go to k. It's k minus one.
- 32:12So this n minus one is has maximum k bits when you do base 2 representation. So you
- 32:19calculate the binary representation. This b 0, b1, b2, bkus 1 are 01.
- 32:26So write this in binary.
- 32:35And uh so now the so some of these b's are zero and the others are one. Uh so the zeros we don't
- 32:46care the once for for example if b2 is one then it means that a to the 2 square has to be calculated
- 32:53and if bkus1 is one then it means that a to 2 to k minus 1 has to be calculated because it
- 32:59contributes to a to the n minus one. Right? So the one the ones the bits here that are on you
- 33:06have to calculate those powers and multiply them this will be the value of a to the n minus one.
- 33:19So,
- 33:24so that so the powering we will do as promised before by this for loop u repeated squaring k
- 33:32times. So we can do that in using the array. So si um will just be s i minus one square mod m.
- 33:50Okay. So that is repeated squaring.
- 33:56So whatever was calculate calculated before that matrix you square mod m and
- 34:04uh once uh so this will create your array s0 s1 and so on.
- 34:16Uh by the way the array is uh correctly ordered. So this basically is a to the 2 to the 0 right
- 34:28so a to the 2 to the 0 is stored in s0 and then in s1 we will be storing a square a to the 2 to
- 34:35the 1 and so on okay so s k minus one I think I don't need to be I don't need k minus one is fine
- 34:48that would give you a to the 2 to the k minus plus one in the end. So the array will be full the s
- 34:55array and which will give you the things that you need. So you need s0 to the b 0
- 35:07multiplied by s1 to the b1
- 35:16s k -1 to the b kus one everything mod m
- 35:26what is this?
- 35:30So what is B now?
- 35:34So by the for loop correctness of the for loop you get that uh S0 is A to the B 0
- 35:43and S1 is
- 35:47s1 to the b1 is a to the 2 b1
- 35:53dot dot 2 to the k -1 b k minus one. So this is nothing but uh a to the n minus one mod m.
- 36:06Okay. So this b that we have calculated by taking a product of essentially array
- 36:12elements which are matrices is the answer that you want. Right? So this answer you output.
- 36:25So you just return b. That's the end. Um and
- 36:32uh that's the matrix mod matrix power mod. But the 1 comma 1 entry
- 36:41that is fn mod m.
- 36:46Okay. So this this matrices top left corner is the uh nth number reduced model 2024. That is the
- 36:57answer. So this is the full pseudo code. Obviously if you convert this into a C program it'll be much
- 37:05longer. Uh but I don't have the energy to give you the C program here. So that you have to do. So as
- 37:13I said I'll make my things simple and your part hard right. So so I'll give you the simple pseudo
- 37:19code then you implement it in C and see whether everything actually works. Uh but the nice thing
- 37:26about writing it first like this is that you can very quickly do a back of the envelope calculation
- 37:32for time complexity. Right? You can see that uh broadly the C program will be doing this
- 37:40and so you can give a good estimation of the time complexity how fast or slow your algorithm is how
- 37:48good or bad. So let us do that. Any questions about this pseudo code? This generally will be
- 37:56the modus operendi. I will give you pseudo codes only at this level, not in further detail. Okay?
- 38:03So you have to understand this when you go home. But if you have quick questions, you can ask
- 38:12any step that you don't understand uh what is the meaning of this uh these arrows and the
- 38:19backslashes in orange and all that. Do you understand everything? Okay. So after this
- 38:25it'll only get harder. So let's uh first make the time complexity statement. So clever fib
- 38:40takes only.
- 38:49So let's put a guess here. So 20k log m was the
- 38:54uh number of steps or time for calculating this last power.
- 39:02Right?
- 39:05So let's write that down. So 20 log m* k. So k means k is login.
- 39:16Right? So this is the time complexity. 20 log m* login is the time complexity to compute
- 39:21uh uh the last thing in the for loop to to reach to basically
- 39:27cover the for loop. It takes you that much time.
- 39:33uh but the the pseudo code has more right the pseudo code also multiplies to calculate B it
- 39:39multiplies those things in the in the array and so so that is kind of another for loop
- 39:45there are two for loops here so this is a for loop but this is also a hidden for loop
- 39:56because what I have written in uh the second line after the for loop loop that you really
- 40:04cannot I mean if you do it in a C program you cannot compute it immediately because
- 40:08you don't know what K is K is some growing something which depends on the input so the
- 40:14only way you know is you have to do another for or while or some loop right so that loop
- 40:19will have K steps so it's basically two for loops each has k iterations so you get double of that
- 40:26so it's double of that so this is the amount of time. So it's around 40 log login time or steps
- 40:45and uh what is the space what is the space requirement here? So array is the biggest uh space
- 40:54consumer here. array has k uh log m data. K is around login. So this is uh around the same space.
- 41:13Okay. But both these values are extremely small. This is a super
- 41:17fast and u very small space requirement. Right? Log you can ignore because it's
- 41:26uh just 11. So both of them are just log n login time and login space. Um
- 41:39yeah so even when n is uh 2 to 64 it'll be very
- 41:44fast. It's a fast algorithm can implement it on your smartphone.
- 41:53So, so this is happening because uh this is a logarithmic scale algorithm.
- 42:05And uh so we have achieved something in
- 42:08logarithmic scale which is much smaller than the linear
- 42:14uh scale which was around n which is uh much smaller than the exponential scale that's 2 to n.
- 42:28Okay.
- 42:32So super fast will be logarithmic that we have achieved. Linear was iterative fib
- 42:37and exponential was recursive fib. So you have seen the whole spectrum. This is the spectrum of
- 42:42uh almost all the problems which you will face in your future. In practice in engineering or in
- 42:50this class all the problems will be in this scale. You will almost never face a I don't
- 42:56think you will ever see a problem which will require more than exponential time.
- 43:00And uh you again it'll be rare to find an find a problem where you will have
- 43:06a faster than logarithmic time. Okay, this will be the scale of
- 43:11uh your uh efficiency and inefficiency. So this is usually written login this is n and this is 2 to
- 43:28n is the input size. Any
- 43:39questions?
- 43:43Okay. Yes. So this discussion till now necessitates uh formalizing time complexity
- 43:55where we don't have to worry about this uh factors of 20 and two and so on. So for that we'll need
- 44:04a proper notation. So let us try to formalize it a bit. So time complexity of a pseudo code
- 44:16or an algorithm.
- 44:27So it is defined as
- 44:35it's the number of uh
- 44:42instructions required
- 44:51in the worst case.
- 45:01Uh it as a function of the input size
- 45:15and usually input size we'll use uh the variable n. Okay, n will be our favorite
- 45:22uh variable to mean input size. Input size is just uh when you write a C program and
- 45:28it takes an input from a file, what is the size of the file? Okay, that is all which
- 45:35uh by uh which n denotes or what we mean by input size. So file is measured in bits or
- 45:41bytes and that is basically what your input size is. and output size will be similar. If
- 45:47your C program takes input from a file and puts input in another file, what is the size of the
- 45:52output file? Right? So these are the input output sizes. Uh everything that we discuss complexity or
- 45:59functions or whatever are functions of uh this n. Uh now the one important thing here is worst case
- 46:09uh if you are not experienced. So generally this is a tendency of students that they write a
- 46:17program only to pass the test cases right. So they they say that for this case my algorithm is very
- 46:24nice and beautiful and fast so I should get full marks right this is the logic. Uh so for that I
- 46:31have underlined worst case analysis. So worst case is against your feelings. So you have to give an
- 46:36algorithm which works for all inputs inputs. Okay. Whether it is given to you in a test case or not.
- 46:43So your the algorithm that you give when you say that it's a fast algorithm it should be fast for
- 46:48all cases which means in the worst case so you should always think adversarially an
- 46:54adversary is giving you an input uh and your algorithm has to behave as you are saying it
- 47:01is expected to behave. Okay. So that is the analysis which we do in this course.
About this transcript
This page contains the full transcript of Lecture 03: Sorting, Searching and Arrays by IIT KANPUR-NPTEL, generated from the public captions YouTube serves with the video. The transcript has 5,168 words across 338 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.