Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Generate the document completely, obtain the finished PDF as bytes, wrap those bytes in io.BytesIO, rewind the stream with seek(0), and upload it with Boto3’s upload_fileobj. This keeps the workflow in memory and lets you set the PDF MIME type with ExtraArgs. If a PDF already exists on disk, use upload_file instead.
The core in-memory upload
A PDF generator and S3 uploader are separate steps. Your generator must finish writing a valid PDF and expose its bytes. The upload function can then treat those bytes as a readable binary file-like object:
from io import BytesIO
import boto3
def upload_pdf_bytes(pdf_bytes: bytes, bucket: str, key: str) -> None:
stream = BytesIO(pdf_bytes)
stream.seek(0)
boto3.client("s3").upload_fileobj(
stream,
bucket,
key,
ExtraArgs={"ContentType": "application/pdf"},
)
upload_fileobj expects a readable object in binary mode that returns bytes. Keeping the BytesIO object alive until the call returns is important, because the managed transfer reads from it during the upload.
Free tools Windows power users keep installed
One-click scans. No signup required.
A complete Python example
The following example shows the handoff from generated bytes to S3. Replace generate_pdf_bytes with the PDF library and layout code used by your application.
#1 Best Overall
from __future__ import annotations
from io import BytesIO
from typing import Callable
import boto3
def upload_pdf_bytes(pdf_bytes: bytes, bucket: str, key: str) -> str:
"""Upload a completed PDF and return its S3 object location."""
if not pdf_bytes:
raise ValueError("pdf_bytes is empty")
if not key.lower().endswith(".pdf"):
raise ValueError("S3 key should end with .pdf")
stream = BytesIO(pdf_bytes)
stream.seek(0)
s3 = boto3.client("s3")
s3.upload_fileobj(
stream,
bucket,
key,
ExtraArgs={"ContentType": "application/pdf"},
)
return f"s3://{bucket}/{key}"
def save_generated_pdf(
generate_pdf_bytes: Callable[[], bytes],
bucket: str,
key: str,
) -> str:
pdf_bytes = generate_pdf_bytes()
return upload_pdf_bytes(pdf_bytes, bucket, key)
# Example integration point:
# def generate_pdf_bytes() -> bytes:
# return your_pdf_library.render(...)
# location = save_generated_pdf(generate_pdf_bytes, "my-bucket", "reports/report-001.pdf")
# print(location) # Only after upload_fileobj succeeds
Return or log the bucket and key only after upload_fileobj completes successfully. The returned s3:// location identifies the object; it is not automatically a browser-download URL.
What the generator must provide
- A finished, valid PDF rather than partially written output.
- A
bytesvalue (or a binary stream that follows the same contract). - A completed generation step before the stream is handed to S3.
Libraries that write to a stream can often write directly to BytesIO; libraries that return bytes can be passed to the function above. The PDF-generation choice depends on document layout and is independent of S3.
Choosing upload_fileobj or upload_file
| Situation | Use | Why |
|---|---|---|
| PDF is already in memory as bytes or a binary stream | upload_fileobj(fileobj, bucket, key) |
Stream-oriented; avoids creating a temporary path. |
| PDF has already been written to disk | upload_file(filename, bucket, key) |
Path-oriented; give Boto3 the local filename. |
The choice is not about PDF format. It is about the source you hand to Boto3. If a local file is acceptable and already exists, upload_file is direct. If your generator naturally returns bytes, upload_fileobj avoids an unnecessary write and read.
Writing the PDF directly into a memory stream
If your PDF library accepts a file-like output, create one BytesIO object, let the library finish writing, rewind it, and upload it. The essential sequence is:
from io import BytesIO
import boto3
def render_and_upload(bucket: str, key: str) -> None:
output = BytesIO()
# Replace this with your generator's stream-writing API.
# pdf_generator.write(output)
output.seek(0)
boto3.client("s3").upload_fileobj(
output,
bucket,
key,
ExtraArgs={"ContentType": "application/pdf"},
)
Do not call the upload while the generator is still writing. A stream’s current position is also significant: after writing, it is commonly at the end, so seek(0) makes the complete document available to the uploader.
Rank #2
Object metadata, progress, and transfer settings
Set the MIME type
Pass supported object settings through ExtraArgs. For a PDF, use ContentType: application/pdf so downstream consumers receive the intended MIME type.
extra_args = {
"ContentType": "application/pdf",
"Metadata": {"document-kind": "invoice"},
}
s3.upload_fileobj(stream, bucket, key, ExtraArgs=extra_args)
Metadata values should describe information your application actually needs; keep the stable object key as the canonical identifier.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11The Boto3 also accepts a transfer An in-memory design is convenient, but the PDF bytes must fit in your process’s available memory while the upload reads them. For large documents or constrained workers, compare these options: Do not claim a universal size threshold: the appropriate approach depends on the PDF, process limits, and deployment environment. Choose a deterministic naming scheme that ends in Do these 3 things before closing this tab: A successful upload does not by itself grant public access. Keep access control and any download or application URL policy separate from the upload function. Cause: the stream position was left at the end, or generation produced no bytes. Fix: verify Cause: invalid or incomplete output from the generator. Fix: finish generation before the upload and validate the produced bytes with the PDF tooling used by your application. Cause: the client credentials or permissions do not allow the requested bucket/key operation. Fix: check the runtime’s AWS credentials, target bucket, region configuration, and permissions; catch and surface the AWS client exception your application expects. Quick wins for a faster PC: Cause: unstable or reused key construction. Fix: centralize key generation, include the intended identifier or version, and decide explicitly whether replacement is allowed. Cause: the callback reports transfer notifications, not a promise that your expected byte count is correct. Fix: compare against the actual byte length and treat callback output as progress telemetry. Cause: the complete PDF is held in memory. Fix: use a temporary file with This isolates PDF-generation errors from S3-transfer errors and makes retries easier to reason about. Recommended Free Tools If the PDF is a screenshot or printout of a web page, ScreenshotNeo can generate the PDF through one API request, so you can upload the response bytes with the same See the ScreenshotNeo documentation for request options. A direct request can return a PDF that you then pass to S3: ScreenshotNeo has 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account. Yes. Keep the generated bytes in memory, wrap them in What’s actually slowing this PC down? Pick the symptom - the matching free tool is one click away. Use it when the PDF already exists at a local path. It is Boto3’s path-oriented helper. No. The file object must be in binary mode and provide bytes. Yes. Pass Do not publish the bucket/key as successful. Catch the AWS client failure your application expects, log useful context without exposing secrets, and apply a retry policy appropriate to the operation. Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.Callback argument can receive transfer progress notifications. A callback can update a progress counter, emit application telemetry, or report status to a job system.class Progress:
def __init__(self, total: int) -> None:
self.total = total
self.seen = 0
def __call__(self, amount: int) -> None:
self.seen += amount
print(f"uploaded {self.seen}/{self.total} bytes")
progress = Progress(len(pdf_bytes))
s3.upload_fileobj(
BytesIO(pdf_bytes),
bucket,
key,
ExtraArgs={"ContentType": "application/pdf"},
Callback=progress,
)
Config. Its managed transfer can use multipart upload and multiple threads when necessary. Choose transfer settings according to your document sizes and runtime constraints rather than assuming every PDF needs custom tuning.Memory use and large PDFs
upload_file when avoiding memory pressure matters more than avoiding disk I/O.Stable keys and application behavior
.pdf, such as reports/<report-id>.pdf or a versioned path. Decide whether retries should overwrite the same key (idempotent replacement) or create a new key. Return the bucket/key only after a successful call, and let the application handle credential, permission, bucket, and network failures.Troubleshooting
The uploaded object is empty
pdf_bytes is non-empty and call stream.seek(0) immediately before uploading.The object is not recognized as a PDF
The upload reports an access or credential failure
The key is wrong or objects overwrite one another
Progress never reaches the expected total
The process runs out of memory
upload_file, reduce simultaneous jobs, or redesign the generator/transfer path around a stream suitable for your workload.Testing the handoff safely
.pdf with ContentType set.Or skip the browser setup
BytesIO pattern. Its cleanup steps accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents such as Claude and Cursor.Best Value
import requests
from io import BytesIO
import boto3
shot = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://stripe.com",
"format": "pdf",
},
timeout=90,
)
shot.raise_for_status()
boto3.client("s3").upload_fileobj(
BytesIO(shot.content),
"my-bucket",
"captures/stripe.pdf",
ExtraArgs={"ContentType": "application/pdf"},
)
FAQ
Can I upload a PDF without writing a temporary file?
BytesIO, rewind with seek(0), and call upload_fileobj.When should I use
upload_file?Does
upload_fileobj accept text streams?Can I attach a PDF content type?
ExtraArgs={"ContentType": "application/pdf"}.What should happen after a failed upload?
Quick Recap

