Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use PyMuPDF to change each page’s CropBox in memory. Set its bottom edge just below the real content, then call Document.tobytes() to obtain new PDF bytes. This changes the visible page area; it does not securely delete objects outside the crop.

The direct in-memory method

PyMuPDF represents the visible page boundary with a CropBox. Open the source bytes, keep the existing left, top, and right edges, replace the bottom edge with a value above the unwanted whitespace, and serialize the document back to bytes. The page’s MediaBox remains the physical boundary while the CropBox controls what viewers display. See the PyMuPDF Page documentation for the API contract.

The operation below assumes you already know the bottom coordinate for each page. It is not an automatic whitespace detector: you must derive and verify that boundary from the document’s actual content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Understand the coordinate system first

PyMuPDF coordinates run downward

PyMuPDF uses an origin near the top-left of an unrotated page, with increasing y values moving downward. Therefore, to remove whitespace at the bottom, choose a new bottom value smaller than the current box.y1. Keep it greater than box.y0 so the rectangle remains non-empty.

The PDF specification commonly describes page coordinates with a bottom-left origin. Do not copy a coordinate calculated in another PDF library into PyMuPDF without converting or checking it. PyMuPDF’s CropBox setter also requires unrotated coordinates. The PyMuPDF FAQ explains why displayed geometry can differ from stored page boxes.

Rotation changes apparent geometry

On a rotated page, page.rect can differ from page.cropbox. Read the page’s box and account for its rotation before selecting the bottom boundary; do not infer an unrotated CropBox coordinate from a screenshot of the rotated page alone.

A complete Python implementation

This function accepts PDF bytes and a bottom coordinate for every page. It validates that each proposed rectangle is finite, non-empty, and inside both the existing page box and the MediaBox before calling set_cropbox(). Check the in-memory opening signature and serialization options against the PyMuPDF version installed in your project, because the documentation’s supported signatures can vary by release.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import math
import pymupdf


def crop_bottom_whitespace(pdf_bytes: bytes, new_bottoms: list[float]) -> bytes:
    """Return PDF bytes with each page's CropBox bottom adjusted."""
    doc = pymupdf.open(stream=pdf_bytes, filetype='pdf')
    try:
        if len(new_bottoms) != len(doc):
            raise ValueError('Provide one bottom coordinate per page')

        for index, page in enumerate(doc):
            box = page.cropbox       # unrotated coordinates
            media = page.mediabox
            bottom = float(new_bottoms[index])

            if not math.isfinite(bottom):
                raise ValueError(f'Page {index + 1}: bottom must be finite')
            if bottom <= box.y0 or bottom > box.y1:
                raise ValueError(
                    f'Page {index + 1}: bottom must be between '
                    f'{box.y0} and {box.y1}'
                )

            candidate = pymupdf.Rect(box.x0, box.y0, box.x1, bottom)
            if (candidate.x0 < media.x0 or candidate.y0 < media.y0 or
                    candidate.x1 > media.x1 or candidate.y1 > media.y1):
                raise ValueError(
                    f'Page {index + 1}: CropBox must stay inside MediaBox'
                )

            page.set_cropbox(candidate)

        return doc.tobytes()
    finally:
        doc.close()


with open('input.pdf', 'rb') as source:
    original_bytes = source.read()

# Replace these example values with boundaries measured for your PDF.
new_bottoms = [720.0]  # one value for a one-page document
cropped_bytes = crop_bottom_whitespace(original_bytes, new_bottoms)

with open('cropped.pdf', 'wb') as output:
    output.write(cropped_bytes)

The PDF is opened and edited in memory; the final two file operations are only a demonstration of how to read and persist the byte strings. In a web service, return cropped_bytes directly in the HTTP response or pass it to the next processing step.

How to choose the new bottom safely

Use a known template boundary

If every page comes from the same report template, measure the first page in the template, retain a small visual margin below the lowest intended content, and reuse that coordinate. Keep one value per page when page layouts differ.

Inspect before changing

Render or view representative pages, note the existing page.cropbox, and compare the proposed bottom with the actual lowest text, image, line, footer, and annotation you intend to retain. A boundary that is too high can clip a footer or signature; one that is too low leaves some whitespace.

Do not treat whitespace removal as automatic detection

The CropBox API only applies the rectangle you provide. It does not discover the lowest visible object. If you need automatic detection, build and test a separate content-bound calculation for your document type, then feed its verified result into set_cropbox(). Keep a safety margin because antialiased strokes, shadows, and annotations can extend beyond an obvious text baseline.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What happens to the original page content?

set_cropbox() changes the visible part of the page. In PyMuPDF’s example, changing the CropBox changes page.rect while page.mediabox remains unchanged. Objects outside the visible rectangle are not automatically proven to be erased; text or graphics may still exist in the file and could remain accessible to extractors or editors.

Use this method when your goal is presentation, printing, or downstream rendering. If your requirement is secure removal of hidden material, treat that as a different workflow requiring a genuine content-removal or redaction process. Do not describe a CropBox edit as sanitization.

Page rotation and mixed documents

  1. Read page.rotation, page.cropbox, and page.mediabox for every page.
  2. Derive the boundary in the page’s unrotated coordinate space, because that is what set_cropbox() expects.
  3. Apply the rectangle, then inspect the resulting page.rect or render a preview. On rotated pages, the displayed width and height can make a correct unrotated rectangle look counterintuitive.
  4. Test portrait, landscape, and mixed-rotation pages separately before processing a large batch.

Never assume that one numeric bottom works for pages with different rotations or source MediaBoxes.

Alternative: pypdf page-box editing

pypdf also documents direct manipulation of page boxes, including cropbox, in its 6.12.2 cropping and transforming guide. The available material does not establish that pypdf is better for this in-memory task, so choose based on the library already used by your pipeline and test the exact installed version. If you combine or merge pages, check compatibility notes: the pypdf guide records merge-behavior changes in versions newer than 3.4.0.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Question PyMuPDF pypdf
Operation discussed here page.set_cropbox(rect) Direct page-box manipulation, including cropbox
Coordinate warning CropBox setter requires unrotated coordinates; PyMuPDF uses a top-left, downward-growing view Follow the coordinate and transformation rules in the installed pypdf documentation
Version detail established by the cited guide Check the installed PyMuPDF documentation for byte-stream and serialization signatures The cited guide is for 6.12.2 and notes merge changes above 3.4.0

Or skip the browser setup

If your workflow also needs screenshots of web pages, ScreenshotNeo provides a one-request website screenshot API; it is separate from editing a PDF’s CropBox. Its cleanup steps accept cookie banners, remove more than 60 known consent platforms plus newsletter popups and chat widgets, and each response reports page and billing status. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed as clean shots. An MCP server lets Claude, Cursor, or another MCP client call take_screenshot, get_page_info, and capture_pdf.

Use the ScreenshotNeo API documentation for authentication and options. The following examples use the supplied endpoint and parameter names:

curl -G 'https://api.screenshotneo.com/v1/shot' 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://stripe.com 
  -o shot.webp
import requests

r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await fs.promises.writeFile('shot.webp', data);

ScreenshotNeo includes 1,000 screenshots each month on the free plan with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account to try it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

“Rectangle is empty” or the setter raises an exception

Your proposed bottom is at or above box.y0, or another edge is invalid. Ensure box.y0 < new_bottom <= box.y1, and reject NaN or infinite values before calling set_cropbox().

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The crop extends outside the page

The CropBox must be completely contained in the MediaBox. Compare every candidate edge with page.mediabox; do not assume the current CropBox and MediaBox have identical coordinates.

The result clips content on rotated pages

You probably measured a displayed, rotated rectangle and supplied it as an unrotated coordinate. Recalculate from the unrotated page boxes, apply the crop, and preview that page before batching.

Whitespace is still visible

Confirm that you changed the intended page and that the proposed bottom is actually smaller than the previous cropbox.y1. Inspect the serialized output in a second viewer; some applications may display a cached preview.

Text outside the crop can still be extracted

That is expected for a visibility crop. CropBox changes the viewing boundary, not a guarantee that underlying objects were deleted. Use a separate redaction or content-removal design when confidentiality is the requirement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The in-memory call fails after a library upgrade

Check the installed PyMuPDF version’s documentation for the supported open(stream=..., filetype=...) form and tobytes() options. Keep a small fixture PDF in your tests and verify page count, rotation, boxes, and rendered output after upgrades.

Performance and reliability considerations

  • Keep the original byte string until validation succeeds so a failed batch can be retried without re-reading a damaged output.
  • Process one page at a time when calculating boundaries, but serialize only after all pages pass validation.
  • Memory use includes the source document and the serialized result. For large PDFs, enforce an application size limit and monitor worker memory rather than assuming cropping is free.
  • Validate every page independently in mixed-size or mixed-rotation documents; a single shared coordinate is safe only for genuinely uniform pages.
  • After serialization, reopen the returned bytes in a test path and verify the CropBox values and a visual sample. This catches coordinate mistakes that syntactically valid PDFs will not.

Practical checklist

  • Decide whether you need a visible crop or secure content removal.
  • Preserve the original PDF bytes.
  • Measure the lowest content edge and choose a margin.
  • Use unrotated PyMuPDF coordinates.
  • Keep the candidate rectangle non-empty and inside the MediaBox.
  • Apply one boundary per page when layouts or rotations differ.
  • Serialize with doc.tobytes(), reopen the result, and inspect representative pages.

Frequently Asked Questions

Can I undo a CropBox change after serialization?

Only if you retain the original PDF bytes or separately record each page’s original CropBox. The safest undo strategy is to keep the untouched input and regenerate the output.

Does cropping change the PDF’s MediaBox dimensions?

No. The CropBox controls the visible area; the MediaBox remains the page’s underlying boundary.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.