Free tools Windows power users keep installed
One-click scans. No signup required.
Short answer: Put a separate payload in a PDF embedded-file stream, then expose it through a file specification. For document-wide files, the specification is normally indexed in the catalog’s EmbeddedFiles name tree; for a paperclip-style attachment, it is referenced by a page annotation. Use /AF Associated Files when the payload has a defined relationship to a page, image, or other object. XMP metadata is for descriptive properties, not a replacement for an attached file.
What “arbitrary data” means in a PDF
A PDF is a collection of indirect objects and streams. Many objects contain bytes, but those bytes are not automatically user-facing files. Text and page content use content streams; fonts, ICC profiles and images use their own object types. A conventional attachment is different: it has an embedded-file stream containing the payload and a file specification containing its name and relationship information. Adobe’s PDF Reference 1.7 documents this model (PDF 1.4 and later): file specifications and embedded file streams.
“Arbitrary” can therefore mean a text file, JSON document, binary blob, source archive or any other bytes you can place in the embedded stream. The PDF viewer may show a filename and offer Save, but it does not need to understand the payload’s format.
Choose the right PDF structure
| Structure | Scope and purpose | Typical reader visibility |
|---|---|---|
| Document-level embedded file | A file specification is indexed in the catalog’s EmbeddedFiles name tree. It applies to the document as a whole. |
Usually listed in an Attachments or paperclip panel. |
| Page attachment annotation | A file specification is attached to a location on one page. This is the familiar paperclip icon beside page content. | Visible when the viewer supports annotations and attachments. |
Associated File (/AF) |
Links a file to a particular PDF object, such as a page, image or structure element, with a machine-readable relationship. | Support varies; intended for interoperable, semantic associations. |
| XMP metadata | Stores small descriptive properties such as title, creator or identifiers in embedded XML metadata. | Shown in document properties or metadata tools, not as a downloadable file. |
| Other streams | Image XObjects, fonts, ICC profiles and content streams hold bytes needed to render the PDF. | Not normally presented as attachments. |
The PDF Association’s overview, Files inside PDF, explains why an inventory of EmbeddedFiles is not universal: 3D, rich-media and other assets can use different structures, and tools may enumerate different sets.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
How do I embed a file in a PDF with Python?
pikepdf 10.15.0 documentation exposes attachments through Pdf.attachments. The mapping accepts bytes for a new payload, and an existing attachment can be read with read_bytes(). The documented model also records the file specification in the catalog’s /AF array when an attachment is added.
Install and add an in-memory payload
python -m pip install pikepdf
import pikepdf
payload = b'{"source":"example","enabled":true}n'
with pikepdf.Pdf.open("input.pdf") as pdf:
pdf.attachments["payload.json"] = payload
pdf.save("output-with-attachment.pdf")
The key is the filename exposed to a viewer; the value is the exact byte sequence to embed. For a binary payload, read it in binary mode and assign those bytes:
with open("archive.zip", "rb") as source:
data = source.read()
with pikepdf.Pdf.open("input.pdf") as pdf:
pdf.attachments["archive.zip"] = data
pdf.save("output.pdf")
Add an on-disk file through an attached file specification
import pikepdf
with pikepdf.Pdf.open("input.pdf") as pdf:
spec = pikepdf.AttachedFileSpec.from_filepath(pdf, "readme.txt")
pdf.attachments["readme.txt"] = spec
pdf.save("output.pdf")
Check the installed pikepdf release before deploying: the examples above follow the 10.15.0 documentation’s API shape, but this article does not claim the snippets were executed here. Preserve the original file until you have opened the result in the viewers and scripts that matter to you.
Rank #2
How do I extract attachments from a PDF?
Iterate over pdf.attachments, call read_bytes(), and write the returned bytes with a chosen output policy. Do not blindly trust an embedded filename: it can contain a path, duplicate another file, or use an unexpected extension.
from pathlib import Path
import pikepdf
source = Path("input.pdf")
out_dir = Path("extracted")
out_dir.mkdir(exist_ok=True)
with pikepdf.Pdf.open(source) as pdf:
for name, attached_file in pdf.attachments.items():
safe_name = Path(name).name or "unnamed.bin"
destination = out_dir / safe_name
# Avoid overwriting an earlier attachment with the same displayed name.
stem, suffix = destination.stem, destination.suffix
counter = 1
while destination.exists():
destination = out_dir / f"{stem}-{counter}{suffix}"
counter += 1
destination.write_bytes(attached_file.read_bytes())
print(f"wrote {destination}")
This extracts the attachments represented by the library’s attachment mapping. It is not a forensic promise that every file-like object in the PDF has been found.
How do I add arbitrary data to a PDF safely?
- Define the consumer. Decide whether people will click an attachment, another program will discover it semantically, or a forensic tool will inspect historical revisions.
- Select the structure. Use document-level
EmbeddedFilesfor a general downloadable file; a page annotation for a page-local paperclip;/AFfor a relationship to a specific object; XMP only for compact descriptive metadata. - Choose a stable filename and media meaning. Keep names simple, document the format and, when useful, include a hash or version inside the payload.
- Write a new revision and verify it. Reopen the output, enumerate attachments, compare byte hashes and test extraction in the target PDF viewers.
- Protect secrets. An embedded file is delivered to anyone who can obtain the PDF. Encrypt the payload separately when confidentiality is required; do not assume ordinary PDF permissions are equivalent to data-at-rest encryption.
Associated Files are especially useful when the payload semantically belongs to a page, image or other object. The PDF Association describes them as a standardized, interoperable, machine-readable way for PDF 2.0 writers to provide information related to a PDF object; the feature was introduced in PDF/A-3 and included in PDF 2.0. See PDF 2.0 Application Note 002.
How do I extract images from a PDF?
An image displayed on a page is commonly an Image XObject, not a conventional attachment. An image extractor may return the encoded stream, decode it to pixels, or reconstruct a file according to the image’s filters and color information. PDF conversion software may have already rescaled or recompressed the source, so the extracted image need not match the original camera or design file byte-for-byte. If the author attached the original image separately, extract that attachment instead of treating the rendered XObject as the source asset.
For a complete inventory, inspect page resources and annotations in addition to the document attachment mapping. The PDF Association notes that attachments, 3D and rich-media assets and other structures can be represented differently; a normal viewer’s attachment panel is not a universal forensic inventory.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Metadata versus an attached file
XMP is embedded structured metadata for descriptive properties. Adobe’s XMP specifications discuss embedding XMP in PDF and reconciling it with non-XMP properties. Use XMP for a title, identifier, authoring information or a small custom property. Use an embedded-file stream when the consumer needs the original bytes of a separate document. Storing a large JSON object as metadata can also create interoperability and size problems; a named attachment is clearer when it must be downloaded or validated independently.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Extraction, revisions and sanitization pitfalls
Duplicate names and unsafe paths
Two attachments can present the same name. Save into a controlled directory, normalize to a basename, avoid overwriting and apply size limits before processing untrusted files.
Passwords, malformed PDFs and encryption
A password-protected or malformed input may fail before attachment enumeration. Obtain authorization and the correct password, then use a parser version appropriate for the file. Treat extracted bytes as untrusted input and scan them before opening.
Incremental updates and hidden historical bytes
PDF incremental updates can leave earlier objects physically present after a later revision marks an item deleted. A forensic question about concealed historical payloads requires revision-aware analysis; the current attachment panel may show only the latest state.
Best Value
Signatures and archival conformance
Adding or removing objects can invalidate a digital signature. PDF/A workflows may impose requirements on embedded files and relationships. pikepdf documents separate remove_attachments and remove_external_access operations and warns that attachments can be integral to signing workflows; its sanitization documentation should be consulted before stripping anything. Work on a copy, record the signature state and validate the final archival profile.
Troubleshooting checklist
- The attachment is not visible: confirm the viewer supports embedded files, inspect the catalog’s name tree, and check that the file was saved after assignment.
- Extraction returns different bytes: you may be extracting a rendered image or recompressed stream rather than an attached source file; compare hashes of the attachment itself.
- Only some files appear: inspect page annotations, Associated Files and rich-media or 3D structures;
EmbeddedFilesalone is not a universal inventory. - The output signature fails: any modification can invalidate a signature; obtain a newly signed output or use a signing workflow designed for the intended incremental update.
- Sanitization removed required content: restore the original and remove only the structure your threat model requires. Attachments may be part of a business or signing process.
- Automation is slow or memory-heavy: process one file at a time, impose attachment-size limits, stream your own input/output handling where supported, and avoid loading untrusted multi-gigabyte payloads without quotas.
Or skip the browser setup
If your workflow also needs screenshots of a source page, ScreenshotNeo provides a website screenshot API and MCP server rather than requiring you to configure a headless browser. One GET request returns PNG, JPEG, WebP or PDF; cookie banners, newsletter popups and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. AI agents can call its MCP tools—take_screenshot, get_page_info and capture_pdf.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for options and response details. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Can a PDF attachment be executed automatically?
Embedding bytes does not make them safe or executable. Treat every extracted payload as untrusted and require an explicit, controlled decision before opening or running it.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteDoes renaming an embedded file change its contents?
No. The displayed filename is part of the file specification; the payload bytes remain separate. Renaming can still affect how a user’s operating system chooses to open the extracted file.
Can I rely on one PDF viewer for interoperability?
No. Viewer support differs, especially for page annotations, Associated Files, rich media and archival requirements. Test with the software used by your recipients and with a programmatic extractor.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

