Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use two libraries for this job: aiohttp downloads the PDF over HTTP, while pypdf reads it and writes only the pages you choose. For a large document, stream response.content to a temporary file instead of calling read(), then convert human page numbers to Python’s zero-based indexes before adding pages to a new PDF.

What the workflow does

The process has four distinct stages:

  1. Open an aiohttp.ClientSession.
  2. Request the PDF and reject non-success HTTP statuses.
  3. Save the response to disk, preferably in chunks for large files.
  4. Open the saved file with pypdf.PdfReader, add selected pages to PdfWriter, and write the result.

aiohttp does not manipulate PDF pages. The pypdf project describes itself as a pure-Python PDF library for splitting, merging, cropping and transforming pages. The two libraries therefore solve separate parts of the problem.

Install the dependencies

Create or activate a virtual environment, then install both packages:

python -m pip install aiohttp pypdf

The cited aiohttp stable client quickstart identifies version 3.14.3, while the versioned pypdf documentation cited for page operations is 6.4.2. APIs can differ between releases, so check the documentation for the versions installed in your application.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Download a PDF with aiohttp

Stream the response to a file

This is the safest general-purpose pattern. It checks the status code, closes the response and session through context managers, and writes 64 KiB chunks without creating one giant response-bytes object.

import asyncio
from pathlib import Path

import aiohttp


async def download_pdf(url: str, destination: Path) -> None:
    timeout = aiohttp.ClientTimeout(total=90)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        async with session.get(url) as response:
            response.raise_for_status()
            with destination.open("wb") as output:
                async for chunk in response.content.iter_chunked(64 * 1024):
                    output.write(chunk)


async def main() -> None:
    await download_pdf(
        "https://example.com/document.pdf",
        Path("input.pdf"),
    )


asyncio.run(main())

The chunk size is an implementation choice, not a PDF requirement. Increase or decrease it after measuring your workload. Streaming prevents the HTTP client from loading the complete response into one Python bytes value; it does not make the later PDF parsing use constant memory.

Read the whole body for a small PDF

For a small, trusted file, this shorter form is convenient:

async with aiohttp.ClientSession() as session:
    async with session.get("https://example.com/document.pdf") as response:
        response.raise_for_status()
        pdf_bytes = await response.read()

with open("input.pdf", "wb") as output:
    output.write(pdf_bytes)

The aiohttp quickstart warns that read(), json() and text() load the whole response in memory. Use the streaming form when file size is unknown or potentially large.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Select pages and create the new PDF

Convert human page numbers to Python indexes

Readers normally count the first page as 1. Python sequences start at 0, so human pages 1, 3 and 4 correspond to indexes 0, 2 and 3:

from pypdf import PdfReader, PdfWriter

reader = PdfReader("input.pdf")
writer = PdfWriter()

for page_index in (0, 2, 3):
    writer.add_page(reader.pages[page_index])

with open("selected-pages.pdf", "wb") as output:
    writer.write(output)

Before indexing, compare every requested index with len(reader.pages). An invalid index raises an error instead of producing a partial document.

Handle an inclusive range

For human pages 2 through 5, subtract one from the start and use an inclusive human end converted to Python’s half-open range:

human_start = 2
human_end = 5

start_index = human_start - 1
end_index_exclusive = human_end

if human_start < 1 or human_end < human_start:
    raise ValueError("Page range is invalid")

reader = PdfReader("input.pdf")
if end_index_exclusive > len(reader.pages):
    raise ValueError(
        f"Document has {len(reader.pages)} pages, "
        f"but the range ends at {human_end}"
    )

writer = PdfWriter()
for index in range(start_index, end_index_exclusive):
    writer.add_page(reader.pages[index])

with open("pages-2-to-5.pdf", "wb") as output:
    writer.write(output)

Here, range(1, 5) visits indexes 1, 2, 3 and 4, which are human pages 2, 3, 4 and 5.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Complete command-line script

The following script accepts a URL, a comma-separated list of human page numbers, or an inclusive range such as 2-5. It downloads to a temporary file, validates selections, and writes the final PDF.

#!/usr/bin/env python3
import argparse
import asyncio
import tempfile
from pathlib import Path

import aiohttp
from pypdf import PdfReader, PdfWriter


async def download_pdf(url: str, destination: Path) -> None:
    timeout = aiohttp.ClientTimeout(total=90)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        async with session.get(url, allow_redirects=True) as response:
            response.raise_for_status()
            with destination.open("wb") as output:
                async for chunk in response.content.iter_chunked(64 * 1024):
                    output.write(chunk)


def parse_pages(specification: str, page_count: int) -> list[int]:
    indexes: list[int] = []
    for item in specification.split(","):
        item = item.strip()
        if not item:
            continue
        if "-" in item:
            left, right = (part.strip() for part in item.split("-", 1))
            start, end = int(left), int(right)
            if start < 1 or end < start:
                raise ValueError(f"Invalid range: {item}")
            human_numbers = range(start, end + 1)
        else:
            number = int(item)
            if number < 1:
                raise ValueError(f"Page numbers start at 1: {item}")
            human_numbers = (number,)

        for human_number in human_numbers:
            index = human_number - 1
            if index >= page_count:
                raise ValueError(
                    f"Page {human_number} is outside a {page_count}-page document"
                )
            indexes.append(index)

    if not indexes:
        raise ValueError("No pages were selected")
    return indexes


def export_pages(source: Path, destination: Path, indexes: list[int]) -> None:
    reader = PdfReader(source)
    writer = PdfWriter()
    for index in indexes:
        writer.add_page(reader.pages[index])
    with destination.open("wb") as output:
        writer.write(output)


async def run(url: str, pages: str, destination: Path) -> None:
    with tempfile.TemporaryDirectory() as directory:
        source = Path(directory) / "source.pdf"
        await download_pdf(url, source)
        reader = PdfReader(source)
        indexes = parse_pages(pages, len(reader.pages))
        writer = PdfWriter()
        for index in indexes:
            writer.add_page(reader.pages[index])
        with destination.open("wb") as output:
            writer.write(output)


def main() -> None:
    parser = argparse.ArgumentParser()
    parser.add_argument("url")
    parser.add_argument("pages", help="For example: 1,3,4 or 2-5")
    parser.add_argument("-o", "--output", type=Path, default=Path("selected-pages.pdf"))
    args = parser.parse_args()
    asyncio.run(run(args.url, args.pages, args.output))


if __name__ == "__main__":
    main()

Run it like this:

python export_pages.py 
  https://example.com/document.pdf 
  1,3,4 
  --output selected-pages.pdf

The script intentionally keeps the original order of the selection. If you request 4,2, the output contains page 4 followed by page 2. Repeated numbers are also retained; remove duplicates in parse_pages if your application requires unique pages.

HTTP details that matter in production

Status, redirects and content

raise_for_status() prevents an HTML error page or an access-denied response from being saved as though it were a PDF. The example permits redirects, which is common for signed download URLs. If your service must not follow redirects, pass allow_redirects=False and handle the 3xx response explicitly.

A URL ending in .pdf is not proof that the response is a PDF. For higher assurance, inspect the response headers and the file signature before handing it to pypdf. A normal PDF begins with the bytes %PDF-, but application-specific validation and limits are still required.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Timeouts, size limits and untrusted URLs

Set a total timeout suitable for your network and document sizes. If URLs come from users, validate the URL scheme and destination according to your application’s security policy, restrict where the downloaded file can be written, and enforce a maximum response size. These are defensive application practices, not guarantees supplied by aiohttp.

Do not trust a remote server to send a correct Content-Length. A streaming loop can count bytes as it writes and stop when an application-defined limit is exceeded. Consider downloading into a private temporary directory and moving the finished output into its final location only after successful parsing and writing.

PDF-specific edge cases

Encrypted files

An encrypted PDF may require a password before pages can be accessed. pypdf can expose encryption state, but the exact handling depends on the file and the installed pypdf version. Detect the condition, obtain credentials through a secure channel, and follow that version’s API rather than assuming every encrypted file can be opened.

Malformed or unusual PDFs

Some files have damaged cross-reference data, unusual page trees or features unsupported by the installed library. Catch parsing and writing exceptions, preserve the original download for diagnosis when policy permits, and report the source URL and failure stage. Do not claim that every syntactically valid-looking PDF will export successfully.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Large documents

Streaming the HTTP response reduces transfer-time memory pressure, but pypdf still has to parse the document and construct the selected pages. Process large jobs in a worker with an appropriate memory limit, avoid downloading the same URL repeatedly, and write each result to a distinct path. A temporary file also allows a failed parse to be retried without repeating the network request.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choosing an implementation pattern

Decision Simple option When to choose the other option
HTTP body handling await response.read() Use chunked streaming when size is unknown or large; aiohttp documents that convenience readers load the whole body.
Page selection Individual zero-based indexes Convert an inclusive human range to a half-open Python range when pages are contiguous.
Temporary storage Directly write a named input file Use a private temporary file when multiple jobs run concurrently or partial files must not appear as finished outputs.
Failure handling Let exceptions reach a top-level caller Catch timeout, HTTP, parsing and writing errors when a service must return structured job status.

Troubleshooting

  • 401 or 403 response: the server requires authentication or rejects the client. Supply permitted headers or credentials, or obtain an authorized URL; do not save the response as a PDF.
  • 404 response: the URL is wrong, expired or redirected elsewhere. Check the URL and redirect behavior.
  • TimeoutError or a stalled download: increase the timeout only when justified, verify connectivity, and enforce a size limit so a slow or unbounded response cannot consume workers indefinitely.
  • IndexError while selecting: the code used a zero-based index that is outside len(reader.pages). Convert human numbers with number - 1 and validate first.
  • “No pages were selected”: the range parser received an empty or whitespace-only specification. Require at least one positive page number.
  • pypdf reports a malformed file: inspect the first bytes and response headers; an HTTP error page is often mistaken for a PDF. If the signature is correct, try the file with the installed version’s repair or compatibility options.
  • Password or encryption error: obtain the document password and handle decryption using the pypdf version installed by your project.
  • Output is unexpectedly large: page resources can be shared or copied during writing. Verify the selected page count and test with the target pypdf release before promising a particular output size.

Or skip the browser setup

If what you actually need is a clean PDF or image of a webpage—not extraction of pages from an existing PDF—ScreenshotNeo can do that with one HTTP call. It accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

This is an alternative for capturing a URL. It does not replace pypdf when the input is an already downloaded PDF and you need to keep selected original pages.

See the ScreenshotNeo API documentation for parameters. The basic cURL request is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account.

FAQ

Does aiohttp extract PDF pages?

No. It performs asynchronous HTTP I/O. Use pypdf or another PDF library for page selection and writing.

Are page numbers passed to pypdf one-based?

No. Accessing reader.pages uses Python’s zero-based indexing. Translate a printed page number before indexing.

Can I export pages without saving the download?

You can keep the response in memory for small files, but a temporary streamed file is generally safer for unknown or large documents.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.