What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use urllib.request.urlopen() for a small, one-off file, or Requests with stream=True for a large download. In both cases, save binary bytes with wb, set a timeout, check the HTTP result, and do not trust a .pdf suffix alone.
Choose the right Python method
Python offers two practical approaches. The standard-library urllib.request module needs no installation and is concise. Requests is a third-party HTTP client with a higher-level interface; Python’s documentation specifically points readers to Requests for that purpose. The Requests documentation reviewed for this guide is for version 2.34.2 and identifies Python 3.10+ support, so check the project documentation if your environment differs.
| Situation | Recommended approach | Reason |
|---|---|---|
| Small, occasional download | urllib.request.urlopen |
Built into Python; short script and no dependency. |
| Large PDF | Requests with stream=True |
Writes chunks incrementally instead of buffering the complete response. |
| Application with sessions, authentication, or shared HTTP settings | Requests | Higher-level request, status, header, cookie, and authentication handling. |
Download a small PDF with the standard library
This is the shortest complete example. urlopen returns a context-manager response, and its body is bytes, so the destination must be opened in binary mode.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →from pathlib import Path
from urllib.request import urlopen
url = "https://example.com/document.pdf"
out = Path("document.pdf")
with urlopen(url, timeout=30) as response:
out.write_bytes(response.read())
print(f"Saved {out} ({out.stat().st_size} bytes)")
The timeout is an example value, not a universal setting. The call buffers the entire response in memory, which is appropriate for a short file but not a good default for a very large PDF.
#1 Best Overall
Choose an explicit destination
Path("document.pdf") writes in the script’s current working directory. Use an absolute path or a directory you create deliberately when the location matters:
from pathlib import Path
from urllib.request import urlopen
url = "https://example.com/document.pdf"
out = Path("downloads") / "report.pdf"
out.parent.mkdir(parents=True, exist_ok=True)
with urlopen(url, timeout=30) as response:
out.write_bytes(response.read())
Decide separately whether an existing file should be overwritten, renamed, or rejected. Python does not impose one policy for your application.
Stream a large PDF with Requests
Requests downloads a response immediately by default. For a large file, pass stream=True and consume it with iter_content(). The context managers below close both the HTTP response and the output file.
from pathlib import Path
import requests
url = "https://example.com/document.pdf"
out = Path("document.pdf")
with requests.get(url, stream=True, timeout=(5, 60)) as response:
response.raise_for_status()
with out.open("wb") as file:
for chunk in response.iter_content(chunk_size=1024 * 64):
if chunk:
file.write(chunk)
print(f"Saved {out} ({out.stat().st_size} bytes)")
The connection timeout and read timeout in this example, and the 64 KiB chunk size, are starting points rather than universal recommendations. Tune them for the server and your workload. raise_for_status() prevents an unsuccessful HTTP response from being accepted as a PDF.
Requests documents this pattern in its Quickstart. Its Advanced Usage documentation also explains why a streamed response should be fully consumed or closed: an unread body can keep the connection unavailable for reuse.
Rank #2
Install Requests (if you use it)
Requests is not part of Python’s standard library. Install it in the environment that runs your script:
python -m pip install requests
Then run the script with the same Python interpreter, for example python download_pdf.py. In a deployed application, pin and manage the dependency according to that project’s normal packaging process.
Handle redirects, status codes, and non-PDF responses
A URL can redirect, require authentication, or return an HTML login or error page. A URL ending in .pdf is only a naming clue; it does not prove that the response body is a PDF.
Requests status handling
Call raise_for_status() before writing bytes, as in the streaming example. If you need a branch instead of an exception, inspect response.status_code and accept the status codes your application defines as successful.
urllib exceptions
urlopen raises URLError for URL and protocol problems. HTTP failures are represented by HTTPError, which is a URLError subclass. Catch them when your program needs a friendly message or a recovery path:
from urllib.error import HTTPError, URLError
from urllib.request import urlopen
try:
with urlopen("https://example.com/document.pdf", timeout=30) as response:
data = response.read()
except HTTPError as exc:
print(f"HTTP error {exc.code}: {exc.reason}")
except URLError as exc:
print(f"URL or network error: {exc.reason}")
else:
with open("document.pdf", "wb") as file:
file.write(data)
Validate when correctness matters
For a workflow that must receive a PDF, check the response status and apply a PDF-aware validation step before handing the file to downstream software. A practical lightweight check is to inspect the beginning of the saved bytes for the PDF signature %PDF-; this is an application-level safeguard, not a substitute for a full PDF parser. Also consider checking the reported content type, while remembering that server headers can be wrong. Do not silently save an HTML access-denied page under a .pdf filename.
Download with authentication or request headers
Some servers require credentials, cookies, an authorization header, or an approved user agent. Supply only credentials you are authorized to use; downloading a protected document is not a way to bypass access controls.
from pathlib import Path
import requests
url = "https://example.com/private/report.pdf"
headers = {"Authorization": "Bearer YOUR_TOKEN"}
out = Path("report.pdf")
with requests.get(url, headers=headers, stream=True, timeout=(5, 60)) as response:
response.raise_for_status()
with out.open("wb") as file:
for chunk in response.iter_content(1024 * 64):
if chunk:
file.write(chunk)
For cookies or several related requests, create a Requests session and configure it once. Keep tokens out of source control and avoid printing them in logs.
URLs that do not end in .pdf
The path may be a download endpoint such as /download?id=123, or it may redirect to a file. Pass the complete URL exactly as supplied, including its query string. HTTP clients can follow ordinary redirects, but your code should still check the final response status and validate the resulting content when a guaranteed PDF is required.
If the endpoint returns an HTML viewer rather than the document bytes, downloading that URL will save HTML. Use the site’s documented download endpoint or an authorized API instead of guessing at hidden URLs.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchMemory, cleanup, and filesystem behavior
- Use binary mode: write with
wborPath.write_bytes(); decoding a PDF as text can corrupt it. - Stream large files: Requests’
iter_content()keeps memory use bounded relative to the complete file. - Always close responses: a
withblock releases the connection even when a streamed body is only partly read. - Prepare the directory: create missing parent directories and choose an overwrite policy explicitly.
- Use timeouts: without one, a stalled server can leave a process waiting indefinitely. Select values based on expected latency and file size.
Command-line and JavaScript equivalents
The Python versions above are the main solution. These equivalent examples are useful when a scheduled job or another service owns the download.
cURL
curl -L --fail --output document.pdf "https://example.com/document.pdf"
-L follows redirects, --fail turns HTTP errors into a failing command, and --output writes bytes to the chosen path.
Node.js
const fs = require('node:fs');
const response = await fetch('https://example.com/document.pdf');
if (!response.ok || !response.body) {
throw new Error(`HTTP ${response.status}`);
}
const file = fs.createWriteStream('document.pdf');
for await (const chunk of response.body) {
file.write(chunk);
}
file.end();
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
404 or another HTTP exception |
The endpoint is wrong, expired, or access is denied. | Confirm the complete URL and permissions; do not save the error body as a PDF. |
| File opens as a web page | The server returned login, consent, or an error HTML document. | Check status and content before accepting the file; authenticate through the documented method. |
| Script hangs | The remote server is slow or no timeout was set. | Set timeout; for Requests, separate connect and read timeouts and adjust them for the workload. |
| Memory usage spikes | The complete response was read into memory. | Use Requests streaming and write each non-empty chunk. |
| Output is corrupt | The response was written as text or was not a PDF at all. | Use binary mode, inspect status, and perform PDF-aware validation. |
| Partial file remains after interruption | The process stopped while writing. | Write to a temporary name, then rename after successful validation if your application needs atomic completion. |
| Certificate or proxy error | The runtime cannot establish the required network connection. | Fix the machine’s trusted certificates or proxy configuration; do not disable TLS verification as a blanket workaround. |
Or skip the browser setup
If your real task is rendering a web page as a PDF rather than fetching an existing PDF file, ScreenshotNeo provides a website capture API. It accepts a URL and can return PNG, JPEG, WebP, or PDF output. Before capture, it can accept cookie and consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing result in headers.
See the ScreenshotNeo API documentation for PDF output and options such as paper size, margins, landscape mode, page ranges, waiting for network idle, custom JavaScript, headers, cookies, and authentication.
Free tools Windows power users keep installed
One-click scans. No signup required.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
These calls show the one-request pattern; configure the response format and PDF settings documented by ScreenshotNeo when you need a PDF rendering. An MCP server also provides take_screenshot, get_page_info, and capture_pdf tools to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Is urllib.request.urlretrieve() still available?
Yes. Python 3.13 documents it as a legacy interface that copies a URL resource to a local file. For new code, urlopen makes timeout, response handling, and cleanup explicit.
Best Value
Should I use one giant timeout for every PDF?
No. A timeout should reflect connection latency and expected transfer time. Requests lets you provide separate connect and read values; the examples use illustrative values that you should adjust.
Can I download a password-protected PDF with these snippets?
The HTTP download step can retrieve an authorized file, but opening an encrypted PDF requires a PDF library and the correct password after the bytes are saved. The HTTP clients do not decrypt document contents.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsFrequently Asked Questions
Is urllib.request.urlretrieve() still available?
Yes. Python 3.13 documents it as a legacy interface. For new code, urlopen gives clearer timeout and response-cleanup control.
Should I use one giant timeout for every PDF?
No. Choose values for the server and file size; Requests supports separate connection and read timeouts.
Can these snippets open an encrypted PDF?
They can download an authorized encrypted file, but decryption requires a PDF library and its password after downloading.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

