Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Generate the document completely, obtain the finished PDF as bytes, wrap those bytes in io.BytesIO, rewind the stream with seek(0), and upload it with Boto3’s upload_fileobj. This keeps the workflow in memory and lets you set the PDF MIME type with ExtraArgs. If a PDF already exists on disk, use upload_file instead.

The core in-memory upload

A PDF generator and S3 uploader are separate steps. Your generator must finish writing a valid PDF and expose its bytes. The upload function can then treat those bytes as a readable binary file-like object:

from io import BytesIO
import boto3


def upload_pdf_bytes(pdf_bytes: bytes, bucket: str, key: str) -> None:
    stream = BytesIO(pdf_bytes)
    stream.seek(0)

    boto3.client("s3").upload_fileobj(
        stream,
        bucket,
        key,
        ExtraArgs={"ContentType": "application/pdf"},
    )

upload_fileobj expects a readable object in binary mode that returns bytes. Keeping the BytesIO object alive until the call returns is important, because the managed transfer reads from it during the upload.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A complete Python example

The following example shows the handoff from generated bytes to S3. Replace generate_pdf_bytes with the PDF library and layout code used by your application.

from __future__ import annotations

from io import BytesIO
from typing import Callable

import boto3


def upload_pdf_bytes(pdf_bytes: bytes, bucket: str, key: str) -> str:
    """Upload a completed PDF and return its S3 object location."""
    if not pdf_bytes:
        raise ValueError("pdf_bytes is empty")
    if not key.lower().endswith(".pdf"):
        raise ValueError("S3 key should end with .pdf")

    stream = BytesIO(pdf_bytes)
    stream.seek(0)

    s3 = boto3.client("s3")
    s3.upload_fileobj(
        stream,
        bucket,
        key,
        ExtraArgs={"ContentType": "application/pdf"},
    )
    return f"s3://{bucket}/{key}"


def save_generated_pdf(
    generate_pdf_bytes: Callable[[], bytes],
    bucket: str,
    key: str,
) -> str:
    pdf_bytes = generate_pdf_bytes()
    return upload_pdf_bytes(pdf_bytes, bucket, key)


# Example integration point:
# def generate_pdf_bytes() -> bytes:
#     return your_pdf_library.render(...)
# location = save_generated_pdf(generate_pdf_bytes, "my-bucket", "reports/report-001.pdf")
# print(location)  # Only after upload_fileobj succeeds

Return or log the bucket and key only after upload_fileobj completes successfully. The returned s3:// location identifies the object; it is not automatically a browser-download URL.

What the generator must provide

  • A finished, valid PDF rather than partially written output.
  • A bytes value (or a binary stream that follows the same contract).
  • A completed generation step before the stream is handed to S3.

Libraries that write to a stream can often write directly to BytesIO; libraries that return bytes can be passed to the function above. The PDF-generation choice depends on document layout and is independent of S3.

Choosing upload_fileobj or upload_file

Situation Use Why
PDF is already in memory as bytes or a binary stream upload_fileobj(fileobj, bucket, key) Stream-oriented; avoids creating a temporary path.
PDF has already been written to disk upload_file(filename, bucket, key) Path-oriented; give Boto3 the local filename.

The choice is not about PDF format. It is about the source you hand to Boto3. If a local file is acceptable and already exists, upload_file is direct. If your generator naturally returns bytes, upload_fileobj avoids an unnecessary write and read.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Writing the PDF directly into a memory stream

If your PDF library accepts a file-like output, create one BytesIO object, let the library finish writing, rewind it, and upload it. The essential sequence is:

from io import BytesIO
import boto3


def render_and_upload(bucket: str, key: str) -> None:
    output = BytesIO()

    # Replace this with your generator's stream-writing API.
    # pdf_generator.write(output)

    output.seek(0)
    boto3.client("s3").upload_fileobj(
        output,
        bucket,
        key,
        ExtraArgs={"ContentType": "application/pdf"},
    )

Do not call the upload while the generator is still writing. A stream’s current position is also significant: after writing, it is commonly at the end, so seek(0) makes the complete document available to the uploader.

Object metadata, progress, and transfer settings

Set the MIME type

Pass supported object settings through ExtraArgs. For a PDF, use ContentType: application/pdf so downstream consumers receive the intended MIME type.

extra_args = {
    "ContentType": "application/pdf",
    "Metadata": {"document-kind": "invoice"},
}
s3.upload_fileobj(stream, bucket, key, ExtraArgs=extra_args)

Metadata values should describe information your application actually needs; keep the stable object key as the canonical identifier.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Report transfer progress

The Callback argument can receive transfer progress notifications. A callback can update a progress counter, emit application telemetry, or report status to a job system.

class Progress:
    def __init__(self, total: int) -> None:
        self.total = total
        self.seen = 0

    def __call__(self, amount: int) -> None:
        self.seen += amount
        print(f"uploaded {self.seen}/{self.total} bytes")

progress = Progress(len(pdf_bytes))
s3.upload_fileobj(
    BytesIO(pdf_bytes),
    bucket,
    key,
    ExtraArgs={"ContentType": "application/pdf"},
    Callback=progress,
)

Boto3 also accepts a transfer Config. Its managed transfer can use multipart upload and multiple threads when necessary. Choose transfer settings according to your document sizes and runtime constraints rather than assuming every PDF needs custom tuning.

Memory use and large PDFs

An in-memory design is convenient, but the PDF bytes must fit in your process’s available memory while the upload reads them. For large documents or constrained workers, compare these options:

  • Generate to a temporary file and call upload_file when avoiding memory pressure matters more than avoiding disk I/O.
  • Use a binary file-like stream and keep it open until the managed transfer returns.
  • Use transfer configuration where multipart behavior or concurrency is relevant to your deployment.
  • Release references to the PDF and stream after success or handled failure so a long-running worker can reclaim memory.

Do not claim a universal size threshold: the appropriate approach depends on the PDF, process limits, and deployment environment.

Stable keys and application behavior

Choose a deterministic naming scheme that ends in .pdf, such as reports/<report-id>.pdf or a versioned path. Decide whether retries should overwrite the same key (idempotent replacement) or create a new key. Return the bucket/key only after a successful call, and let the application handle credential, permission, bucket, and network failures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A successful upload does not by itself grant public access. Keep access control and any download or application URL policy separate from the upload function.

Troubleshooting

The uploaded object is empty

Cause: the stream position was left at the end, or generation produced no bytes. Fix: verify pdf_bytes is non-empty and call stream.seek(0) immediately before uploading.

The object is not recognized as a PDF

Cause: invalid or incomplete output from the generator. Fix: finish generation before the upload and validate the produced bytes with the PDF tooling used by your application.

The upload reports an access or credential failure

Cause: the client credentials or permissions do not allow the requested bucket/key operation. Fix: check the runtime’s AWS credentials, target bucket, region configuration, and permissions; catch and surface the AWS client exception your application expects.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The key is wrong or objects overwrite one another

Cause: unstable or reused key construction. Fix: centralize key generation, include the intended identifier or version, and decide explicitly whether replacement is allowed.

Progress never reaches the expected total

Cause: the callback reports transfer notifications, not a promise that your expected byte count is correct. Fix: compare against the actual byte length and treat callback output as progress telemetry.

The process runs out of memory

Cause: the complete PDF is held in memory. Fix: use a temporary file with upload_file, reduce simultaneous jobs, or redesign the generator/transfer path around a stream suitable for your workload.

Testing the handoff safely

  1. Generate a known-small PDF and assert that the result is non-empty bytes.
  2. Upload to a test key ending in .pdf with ContentType set.
  3. Verify that the upload call succeeds before publishing the returned location to application code.
  4. Exercise invalid credentials, an unavailable bucket, a duplicate key, and a network interruption in your application’s error path.
  5. Test large documents separately if your production workload can exceed worker memory.

This isolates PDF-generation errors from S3-transfer errors and makes retries easier to reason about.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If the PDF is a screenshot or printout of a web page, ScreenshotNeo can generate the PDF through one API request, so you can upload the response bytes with the same BytesIO pattern. Its cleanup steps accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents such as Claude and Cursor.

See the ScreenshotNeo documentation for request options. A direct request can return a PDF that you then pass to S3:

import requests
from io import BytesIO
import boto3

shot = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={
        "access_key": "YOUR_API_KEY",
        "url": "https://stripe.com",
        "format": "pdf",
    },
    timeout=90,
)
shot.raise_for_status()

boto3.client("s3").upload_fileobj(
    BytesIO(shot.content),
    "my-bucket",
    "captures/stripe.pdf",
    ExtraArgs={"ContentType": "application/pdf"},
)

ScreenshotNeo has 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Can I upload a PDF without writing a temporary file?

Yes. Keep the generated bytes in memory, wrap them in BytesIO, rewind with seek(0), and call upload_fileobj.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When should I use upload_file?

Use it when the PDF already exists at a local path. It is Boto3’s path-oriented helper.

Does upload_fileobj accept text streams?

No. The file object must be in binary mode and provide bytes.

Can I attach a PDF content type?

Yes. Pass ExtraArgs={"ContentType": "application/pdf"}.

What should happen after a failed upload?

Do not publish the bucket/key as successful. Catch the AWS client failure your application expects, log useful context without exposing secrets, and apply a retry policy appropriate to the operation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.