Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Paginate through a REST API

Python · Python for Data Pipelines

Paginate through a REST API

Mediumpython-42
scenarioapipaginationgeneratorscursor

Question

How do you pull all records from a paginated API in Python without losing or duplicating data?

Solution

The safe way is to use cursor pagination if the API offers it, wrap the fetching in a generator that yields one page at a time, save your progress as you go, and deduplicate by id when you load. Each piece guards against a different failure.

Offset versus cursor

With offset pagination (?page=3&limit=100), the server counts rows from the top each time. If a record is added or deleted while you are reading, the rows shift. You then see one record twice, or skip one without any error. With cursor pagination (?after=abc123), the server returns a token that marks a position, so the next page continues from exactly there, however the data changes. Prefer cursors when available.

A generator keeps memory flat

import requests

def fetch_all(url, token, cursor=None):
    while True:
        params = {"limit": 500}
        if cursor:
            params["cursor"] = cursor
        resp = requests.get(url, params=params,
                            headers={"Authorization": f"Bearer {token}"}, timeout=30)
        resp.raise_for_status()
        body = resp.json()
        yield body["items"], body.get("next_cursor")
        cursor = body.get("next_cursor")
        if not cursor:
            break

The caller processes one page at a time, writes it out, and never holds the full result. The stop condition is the missing next cursor. Do not stop when a page is "short", because some APIs return fewer items for reasons that have nothing to do with being the last page.

Resume after a crash

Save the cursor after each page is safely written (a small state file or a control table). If the job dies at page 4,000, the next run starts there and not at page 1. Write the page and the cursor together, in that order, so a crash can only cause a repeat of one page, never a gap.

Deduplicate downstream

Even with careful code, an overlap or a retry can deliver a record twice. Load by the record's id, with an upsert, or deduplicate in the next step. That makes the whole thing safe to rerun.

Extras

Set a timeout on every request, and cap the number of pages per run, so a bug that returns the same cursor forever cannot loop endlessly. Keep the raw responses if the API is flaky, so you can reprocess without calling again.

🎯 Put this concept into practice

Solidify this answer with real hands-on interview drills in the browser studio.

Open related drill →
PreviousNext