API Pagination - Offset and Cursor Pagination Explained - Backend Engineering — Transcript
Full transcript
- 0:00Hey, what's going on? In this lesson,
- 0:01we're going to talk about pageionation
- 0:02for APIs. We're going to talk about a
- 0:04couple of different approaches, why you
- 0:06might want to page your data at all, and
- 0:09then what it's going to look like from
- 0:11the client's perspective, the different
- 0:13approaches to consuming a pageionated
- 0:15API. Now, before we get started, I want
- 0:17to mention I have a link to the notes,
- 0:19which will have all the information we
- 0:20talk about here with references,
- 0:22resources, as well as examples. So, you
- 0:25should definitely check that out. We'll
- 0:26have a link down below. This is also
- 0:28part of a larger API and backend
- 0:29playlist. So I'll have a link for that
- 0:31playlist too. So first off, what is
- 0:32pageionation
- 0:37and why would you want to do it? So
- 0:40let's say we have 10,000 elements
- 0:44and let's say we want to retrieve this
- 0:47data from an API. The default behavior
- 0:50may be to select all 10,000 of those
- 0:52elements and send that to the user.
- 0:54Well, what this is going to do is one,
- 0:57put a lot of effort on the database to
- 1:00retrieve all of that data, and two, when
- 1:02we send that over the line to the
- 1:04client,
- 1:07we have to send a lot of data.
- 1:13This is ultimately going to slow down
- 1:14the response from the API because the
- 1:17amount of data is so large. So, what we
- 1:19could do instead is we could split this
- 1:21up into pages. Let's say groups of 100.
- 1:24We send that to the client and most of
- 1:26the time that's going to be enough as
- 1:29the times you go from page one to page
- 1:32two to page three
- 1:35is quite small actually. So most of the
- 1:40visits will be on these first pages.
- 1:43I mean just think about it with a Google
- 1:45search. How often do you go to page two
- 1:47of Google?
- 1:51you know, maybe that's 10% of the time
- 1:53or something like that. And then from
- 1:54there, 10% of the time, you'll go to
- 1:56page three. So almost all of that data
- 1:58is not even going to be needed. Now, you
- 2:00have to think about how the client might
- 2:02use this pageionated API because if they
- 2:05do need all of this data for some
- 2:07reason, now they're going to have to
- 2:08make tons of requests grabbing the data
- 2:11100 at a time. So, we also don't want to
- 2:13up the request count too much.
- 2:17So, that's one of the potential
- 2:18tradeoffs. If you reduce the page size,
- 2:21you're going to get more requests to the
- 2:23server for more data. So you have to
- 2:25think of the ideal page size.
- 2:28How many elements do you want to return?
- 2:30And is this something that you can
- 2:32configure from the client?
- 2:35Those are things you need to consider
- 2:38when you're designing your API. And this
- 2:40might be something that's not universal
- 2:42across all of your different resources
- 2:44in your API. There might be certain ones
- 2:46where you adjust the default page size
- 2:50for that use case. If in many scenarios
- 2:53the user is going to work through a lot
- 2:55of this data, you can just create a
- 2:56custom endpoint
- 2:59appropriate for whatever the use case
- 3:02is.
- 3:04I mean, you could think of an example
- 3:05like this where we might have some
- 3:07campaigns or something and you could
- 3:09page it or you could just have a custom
- 3:10endpoint like slashall. If you know you
- 3:13have a web page where you need to
- 3:15display all of those values, that might
- 3:17be okay for some things, but for others,
- 3:20it's just not going to work. I mean,
- 3:22think of like YouTube. They couldn't
- 3:24just grab all of the comments in a
- 3:26single query. So most of the time we're
- 3:28going to concern oursel with paging and
- 3:31then the appropriate page size for the
- 3:33client. So we're going to talk about
- 3:35some of those things now. So first let's
- 3:37talk about the API structure for a
- 3:39pageionated API endpoint. And this is
- 3:42one of the reasons I recommended nesting
- 3:44your data in some attribute. Let's just
- 3:47say that attribute is data.
- 3:50And then we have our data here inside of
- 3:54an array. And having it in this kind of
- 3:57structure allows us to easily add in
- 3:59other attributes here that can be used
- 4:01for paging through the data. So then the
- 4:03client knows if they want to access the
- 4:05data, they just go into the data
- 4:07attribute. If they want to access paging
- 4:10data, they can grab whatever else such
- 4:12as
- 4:14next. So what kind of attributes would
- 4:16we put in here? This is going to depend
- 4:18on your approach to pageionation. We're
- 4:20going to talk about a couple of
- 4:21different approaches, but just to get us
- 4:23started, I'll show you an example.
- 4:25We may have something like the count
- 4:28which is the total number of elements.
- 4:33We may have for ease to the client a
- 4:38next attribute which could be a URL to
- 4:42the next page and then a similar thing
- 4:44for the previous page. So the consumer
- 4:48of this API
- 4:54can navigate very easily just grabbing
- 4:57those attributes. But we can customize
- 5:00this behavior. We can do whatever
- 5:01structure we want as long as the client
- 5:04knows how to take this information and
- 5:06go to the next page. So let's now talk
- 5:09about the types of pageionation.
- 5:15And now we're going to cover two main
- 5:17classifications with the first
- 5:19classification having kind of two
- 5:21different approaches. So we will talk
- 5:24about offset based
- 5:31and we will talk about cursorbased.
- 5:36These are two different approaches with
- 5:38different pros and cons. Sometimes
- 5:40you'll have a business decision that
- 5:41will force a certain one of these. Other
- 5:44times you can choose between the two. So
- 5:46we'll talk about when you have to use a
- 5:49certain one of these. So first let's
- 5:50talk about offsetbased pageionation.
- 5:53This will use a page attribute or it
- 5:57will use an offset and a limit.
- 6:03These are basically two different
- 6:05approaches or two ways of describing the
- 6:07same capability. So we're going to first
- 6:10take a look at just working with a page
- 6:12number. Then we'll talk about how to
- 6:14customize it a bit with an offset and a
- 6:16limit. After we go over offset based, we
- 6:18will then circle back and look at
- 6:20cursorbased pageionation. So let's say
- 6:22we have an endpoint slash comments
- 6:25and then you use a URL parameter
- 6:29such as page and then specify a value
- 6:32such as three. This request will be sent
- 6:34to our backend and the backend is going
- 6:37to have some attribute page size.
- 6:42Let's just say it's a default of 20.
- 6:45Now, how you split up the data is
- 6:47totally up to you in this situation. I'm
- 6:49going to do it by ID, but you could sort
- 6:52it in any way. IDs are very easy to
- 6:54think about cuz they're numbered, but
- 6:56you could sort by any column. So, let's
- 6:58take a look at what page one might look
- 7:00like. That would be IDs 1 through 20.
- 7:06Page two would be 21 through 40.
- 7:12and then page three would be 41 through
- 7:1660. So by specifying page three, we
- 7:19would retrieve this data. And then
- 7:22inside of our API endpoint, we would
- 7:23have a next attribute pointing to the
- 7:25next page. And it would just be the same
- 7:27URL but with a page equal to four. So
- 7:29really simple for the URL navigation. By
- 7:32far the easiest for the client to
- 7:34understand. And then a previous would
- 7:36just be page two. So the question then
- 7:37is how do we go from a number here to an
- 7:39actual database query?
- 7:42Let's write up what this query might
- 7:44look like. So if the page is three, we
- 7:47could have select
- 7:50everything
- 7:52from whatever table.
- 7:55Let's say comments order by
- 7:59ID or comment ID. And then we can pass
- 8:02to the database a limit.
- 8:05This is going to be the set page size.
- 8:08So 20 how many elements per page and
- 8:11then we will provide an offset
- 8:16which is how many rows to skip. So I
- 8:18think we mentioned page one was 1
- 8:21through 20 then page two was 21 through
- 8:2640 and then page three was 41 through
- 8:2960. So if we wanted to jump to page
- 8:32three of our data we would use an offset
- 8:35of 40. Now the important thing to
- 8:37understand here is that this is going to
- 8:39start with the 41st row.
- 8:43So this does not have knowledge of the
- 8:46actual ID value. This is going off of
- 8:49row counts. It just so happens that we
- 8:51have a perfectly sequential number from
- 8:53one all the way up through 40. So that's
- 8:55our first approach really is just taking
- 8:58a page from the client and converting it
- 9:00to some query. But you'll see here we
- 9:02have this limit and offset. We're really
- 9:04just translating that page number into
- 9:06values here. We could just expose these
- 9:08two things to the client and allow them
- 9:10to provide arbitrary values for these
- 9:12and that'll be very similar to the page
- 9:14structure. It's just a little bit more
- 9:16exposed to the client. Even before
- 9:19jumping to that approach though, you can
- 9:20give customization to the client. So for
- 9:22example, you could allow them to provide
- 9:26a page size as a URL parameter and then
- 9:29that could change the query to something
- 9:32else using this value for the limit
- 9:37and also the math to adjust the offset.
- 9:42So you can think of the offset
- 9:46as just the page minus one times the
- 9:51page size.
- 9:54So for example to get the value 40 from
- 9:56our previous example we had page 3. 3 -
- 9:591 is 2 and then 2 * the page size. By
- 10:04default it was 20. So that gives us 40.
- 10:08So it's a really simple operation to go
- 10:10from a page number to the appropriate
- 10:13offset. So if we wanted to allow these
- 10:15to be provided from the user, we might
- 10:17get something like slash comments
- 10:22and then a query parameter limit being
- 10:26equal to say 30 and an offset of
- 10:32120. And now we don't have specific page
- 10:35numbers. Instead, we just have these
- 10:36values which are then translated to
- 10:39grabbing a certain number of rows from
- 10:40the database. And this is really simple.
- 10:42We just pass this directly to the
- 10:45database. Limit 30
- 10:50offset
- 10:52120. Now, be sure to do any kind of
- 10:55parameterization or filtering so we
- 10:57don't get any SQL injection attacks. But
- 11:00that's basically what we're doing. We
- 11:01just substitute those values in
- 11:03directly. then the database just grabs
- 11:05the appropriate rows. So the page ID and
- 11:08the limit and the offset ID are the same
- 11:11thing behind the scenes. It's just
- 11:13changing how the user interacts with the
- 11:16API. The next attribute
- 11:19would look something like this. Comments
- 11:24limit 30
- 11:26and
- 11:30offset 150.
- 11:33So basically just adding the limit to
- 11:36that previous offset to get the next URL
- 11:39and then the client can easily click
- 11:41this if they're within an API viewer or
- 11:44just grab that in code and navigate to
- 11:47that path to get the next page of data.
- 11:49Same idea with the previous you would
- 11:51just subtract instead of add. So I kind
- 11:53of like this approach because it gives
- 11:54control to the client
- 11:57but you can still set boundaries on the
- 11:59back end. So you could say there's a max
- 12:02limit so they don't just grab all of the
- 12:04data. So if you went in here and put
- 12:063,000, maybe it would default to, you
- 12:08know, a max of 150 or something. But it
- 12:11does give the client more control and
- 12:12you can have a drop down that says
- 12:15elements per page
- 12:18and maybe have 20, 50, 100 so that they
- 12:24can choose the appropriate one for
- 12:25whatever they're trying to do. So what
- 12:27are some of the big downsides of this?
- 12:29Well, the big one is that this offset
- 12:34is rowbased. It has no concept of the
- 12:37data or which element it is at. It's
- 12:40just a number that represents how far
- 12:42off we are from the start of that data
- 12:45set. So I should say something like row
- 12:47number based. What this means is if the
- 12:50data changes as you are navigating
- 12:53through the pages, you could get gaps or
- 12:56repeating values that show up. Let's
- 12:58just show a very simple example of this.
- 13:04So here we have 10 elements and let's
- 13:07say we're getting three at a time. So
- 13:08we're looking at these three. The offset
- 13:11here is three. And now let's say we want
- 13:12to go to the next page but some data was
- 13:14deleted. So let's say we have 1 2 5
- 13:206 7 8 9 and 10. We go from an offset of
- 13:24three to an offset of six to go to the
- 13:26next page. So 1 2 3 4 five six and we
- 13:31start here and we go all the way to the
- 13:34end because we're out of data or you
- 13:35could think of this as grabbing 11 as
- 13:37well. So we basically just skipped over
- 13:40seven and 8 because we're using an
- 13:42offset. So basically anytime we add or
- 13:45remove data earlier on in our data set
- 13:48that's going to shift where we start.
- 13:50And if we're having that happen as we're
- 13:52navigating through the pages that's a
- 13:54pretty disruptive experience for the
- 13:55user. So now I want to talk about
- 13:57cursorbased
- 13:59and you can think of this as using a
- 14:01cursor or pointer to the last record
- 14:04scene and this will pretty much ignore
- 14:07any deletes or new data because it
- 14:10doesn't use an offset. You can think of
- 14:13it as just identifying a specific row
- 14:16and then going from there. So going back
- 14:18to that previous example. So let's say
- 14:20the first page had 1 2 3 and then we go
- 14:22to the next page and we grab four five
- 14:25and six
- 14:27and then some data is deleted. So it
- 14:29goes 1 2 5 6 7 8 9 and 10. Well the
- 14:35cursor is basically going to say hey
- 14:37this was our last scene.
- 14:41So this is where we're going to continue
- 14:42from and then we'll go to seven. So
- 14:45we're able to grab that next set of
- 14:46data. So in this situation, we are
- 14:49immune to data being added or removed.
- 14:52But there are a few gotchas here that
- 14:54you definitely need to be aware of. I'm
- 14:55going to talk about some of the pros and
- 14:56cons and when you can and cannot use
- 14:58cursorbased pageionation, but first I
- 15:00want to show you what the actual API
- 15:02response might look like. So to navigate
- 15:04to a page, you might see something like
- 15:06this
- 15:08comments and then we'll start with this.
- 15:12we'll say
- 15:15after id and then some value
- 15:19and then basically the backend can use
- 15:22that value to know where to start and
- 15:24then if we had an API response with a
- 15:26next attribute
- 15:28so it might be something like comments
- 15:31after a ID 1 2 3 7 8 or whatever it may
- 15:37be so in this situation we're still
- 15:39using a transparent URL where the value
- 15:42can be seen and interpreted by the
- 15:44client. In other words, they know what
- 15:46ID we're starting at. And generally,
- 15:49this is considered bad practice. We
- 15:51don't want to expose too much here. And
- 15:54it's also very limiting. Uh so here's
- 15:57some of the reasons. Risk of bad
- 15:58exposure, sharing too much information
- 16:01about the database structure and what
- 16:03rows exist. Limited use. In this case,
- 16:05it's tied to an ID. It's also confusing
- 16:09for the direction because we're saying
- 16:11after which assumes, you know, we're
- 16:13only going forward.
- 16:15So, usually all of this stuff is just
- 16:18hidden from the client and instead we're
- 16:22given a token.
- 16:24This may be called a continuation token
- 16:26or something of that nature,
- 16:31which basically encodes this
- 16:33information. you might be able to decode
- 16:36it. So, it's not life or death. It's not
- 16:39like a secret key or something. But
- 16:42basically, instead of having after ID 1
- 16:442 3 4 5, we might have something a
- 16:47little bit more vague, such as
- 16:53continuation
- 16:55and then some value.
- 16:59And this doesn't tie us into a certain
- 17:02column or ordering or anything like
- 17:04that. We can change the behind the
- 17:06scenes implementation if needed, but the
- 17:08way to interact with the API stays the
- 17:10same. So it's pretty common to see
- 17:12continuation tokens, but the same
- 17:15limitations and behaviors apply. So
- 17:17let's talk about some of the limitations
- 17:19of this structure. But before we erase
- 17:21all of this, you can imagine this value
- 17:25being used in an SQL statement where
- 17:30ID is greater than 1 2 3 7 8. So it goes
- 17:36by the actual ID value and continues
- 17:38from there. With the continuation token,
- 17:40it'll be very important for you to set
- 17:42the next attribute. Since the client is
- 17:47probably not going to be able to type in
- 17:49arbitrary values or calculate it, it'll
- 17:51just provide what value to use next for
- 17:54the next set of data. So, it's a little
- 17:56bit more complicated in setup and
- 17:58understanding from the client, but it
- 17:59does have some of those benefits. When
- 18:01data is added or removed, you're not
- 18:03going to have that shift and then you
- 18:04see some repeating data or skip data.
- 18:06So, let's talk about some of the
- 18:07limitations, some of the reasons why you
- 18:09would not want to do this. So, it's
- 18:11going to be limited to certain columns.
- 18:14Basically, you can think of many of the
- 18:15same attributes of a good primary key.
- 18:18It's going to be unique.
- 18:20Since we need to be able to uniquely
- 18:22identify it, it will have to be not
- 18:25null,
- 18:26unchanging,
- 18:28and we will want it to be indexed
- 18:32to get the most performance here. So,
- 18:35two good examples would be an ID or a
- 18:38creation timestamp
- 18:45because this isn't going to be null.
- 18:47It's something you could index by. It's
- 18:49unique. Most likely if you're not
- 18:51creating multiple at the same exact
- 18:52time, it's unchanging and you can sort
- 18:55by that very easily. So it could be
- 18:57something like the latest comments. So
- 18:59this is something to keep in mind
- 19:00because with the previous pageionation,
- 19:02it's very nice to use for arbitrary
- 19:05filtering and querying. So, if you have
- 19:07a big data set, I think I used in an
- 19:09earlier video the example of a car
- 19:11database and you want to sort and filter
- 19:14by make, model, price, all of these
- 19:17different attributes, that's where a
- 19:19paging system could be quite nice
- 19:21because we're not limited to some of
- 19:23these things with continuation tokens.
- 19:26And this is a big one. Depending on if
- 19:28you need this or not, you're not going
- 19:30to be able to do arbitrary page jumps.
- 19:33It's pretty much just next
- 19:36and previous.
- 19:39Visually, you can think of this almost
- 19:41like a linked list. So, you might have
- 19:43some page of data you're looking at and
- 19:46altogether you think of this as the
- 19:48page, the previous pointer and the next
- 19:52pointer. Meaning, you can't just jump
- 19:54seven pages forward like you could with
- 19:57the previous example. So what that means
- 20:00is client side
- 20:03if you want your current page 17 and
- 20:06then you have the option to click 16 and
- 20:1015 and then we might have 18 and 19.
- 20:13This is not ideal
- 20:17for this type of pageionation. Now it's
- 20:21technically possible you could
- 20:24manipulate this. So, for example, you
- 20:26might have page 17 and then if the user
- 20:29clicked page 19, you could just go to
- 20:31the next pointer twice, but it's just
- 20:34not natural. It's like forcing it. So,
- 20:36if you want the user to be able to say,
- 20:39"Hey, yeah, I want page 36," then you're
- 20:42going to want to go with the offset
- 20:44approach. Or if you want this to be
- 20:46sorted by some arbitrary column,
- 20:54then you will want to go with the offset
- 20:56approach. So client side, the way you
- 20:58would use this is at the bottom of the
- 21:00page, you could have a load more
- 21:06or you can have an infinite scroll,
- 21:08meaning once you go to the bottom,
- 21:12it will just automatically load the next
- 21:15page. This is the perfect way to consume
- 21:17a cursor-based pageenated API. Now, one
- 21:20last thing to mention with the
- 21:21cursorbased approach is that it's the
- 21:23ideal choice for large data sets.
- 21:29So when you're using an offset, the
- 21:31database still has to do some work to
- 21:34jump that offset. You can think of it as
- 21:37just burning through all of those rows
- 21:39to get to where you're trying to go. But
- 21:41if you're working with a specific ID,
- 21:43the database is able to just jump to the
- 21:45proper location. So this is a much more
- 21:48friendly option for very large data sets
- 21:51because in theory every time you go an
- 21:53additional page on a regular offsetbased
- 21:56pageionation based API the query is
- 21:59going to take just a little bit longer
- 22:01and a little bit longer until you're way
- 22:03deep into the records and you have a
- 22:06very large offset and you're wasting
- 22:07compute power by not using the database
- 22:09in the most ideal way. This usually
- 22:12isn't a problem for medium to, you know,
- 22:15regular size data sets, but it is
- 22:16something to consider if you are working
- 22:18with large data. So, let's take a look
- 22:20at these approaches from the client's
- 22:21point of view. And I want you to think
- 22:23of the options of having multiple pages
- 22:25like we talked about earlier. If you're
- 22:28going to go this route, you pretty much
- 22:29want to use that offset because the user
- 22:33can provide an arbitrary page. likely in
- 22:35the URL you're going to see something
- 22:37like page or offset
- 22:41that you can just provide a value for
- 22:43and clicking these buttons just makes it
- 22:45easier for the option of a load more or
- 22:48an infinite scroll. There's one thing I
- 22:51wanted to call out.
- 22:54I think you can do this approach with
- 22:57either of these because basically with
- 22:59an offset base approach you could just
- 23:00load the next page and append that data.
- 23:03But the problem is if you have a current
- 23:06viewed data set. So this is the stuff on
- 23:08the page. It's already loaded in memory.
- 23:12What happens if we click load more but
- 23:15before we clicked load more data was
- 23:18added or removed that affects that
- 23:20offset? This could cause duplicate
- 23:22values in the view set. So let's go
- 23:26through a quick example. We're going to
- 23:28list some comments and let's say we are
- 23:32looking at these three. So these are
- 23:35added to our comments state and then we
- 23:38add some data that affects the data
- 23:40that's being retrieved.
- 23:43This could shift the pages. So we were
- 23:46originally looking at three. When we
- 23:48want to grab the next set of data to
- 23:50append it to our local state, it's going
- 23:53to grab from here. So you can see it
- 23:56crosses over this love it example twice.
- 24:01Additionally, this new comment is
- 24:03skipped over, which isn't probably the
- 24:05end of the world, but our current state
- 24:08is going to look like, "Hey, great
- 24:11videos.
- 24:15Love it."
- 24:18A second instance of love it.
- 24:22And then not bad.
- 24:25This would be avoided with a
- 24:27continuation token because that wouldn't
- 24:29have shifted with the insert of a new
- 24:31comment. So basically what you need to
- 24:33do is if you're appending to a local
- 24:35state, you need to dduplicate,
- 24:38get rid of any duplicate data so the
- 24:40visual doesn't show repeating values.
- 24:42Now this is all just talking about the
- 24:44local state. If you do a refresh, you're
- 24:46not going to have duplicate data. I'm
- 24:48talking about adjusting the state on the
- 24:50page as we go with data being changed on
- 24:53the back end. With a continuation token,
- 24:55this will be a seamless experience. You
- 24:57can just extend the state every time you
- 24:59load more. You're never going to have to
- 25:00worry about the duplicate data showing
- 25:02up. And then data such as this one, this
- 25:05great stuff that was added later,
- 25:08you can just ignore that as you extend
- 25:10the page. And then if you do a refresh,
- 25:13it'll be retrieved fresh. That's all I
- 25:14have in this lesson which was your
- 25:16introduction to pageionation and some of
- 25:18the different approaches. The next
- 25:20logical thing for you to do would be to
- 25:22implement this in some library. So build
- 25:24a backend, create multiple pages and
- 25:26then if you want you can learn to
- 25:28consume that on the client end as well.
- 25:30Thank you so much for watching.
- 25:31Hopefully these API videos have been
- 25:33really helpful. If so, please let me
- 25:34know. Or if there's some other topics
- 25:36you'd like me to discuss, leave a
- 25:37comment so I can adjust what type of
- 25:39material to focus on. Thank you so much
- 25:41for watching and I will see you in the
- 25:43next one. Peace out.
About this transcript
This page contains the full transcript of API Pagination - Offset and Cursor Pagination Explained - Backend Engineering by Caleb Curry, generated from the public captions YouTube serves with the video. The transcript has 4,182 words across 597 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.