Free tools Windows power users keep installed
One-click scans. No signup required.
The release history matching this topic is Crawl4AI’s: version 0.9.4 is listed as the latest release on 23 September 2026, while version 0.9.3 contains the concentrated PDF-path security work and version 0.9.0 introduced the secure-by-default Docker server. The practical lesson is architectural: browser request controls do not automatically protect a separate PDF downloader, so the PDF path needs its own redirect, network, resource, and output controls.
This article treats Crawl4AI as the intended project because that is what the documented release themes identify. If you meant another scraper, verify its version and security model before applying these settings.
What changed, at a glance
Crawl4AI’s recent changes address two different trust boundaries. The PDF fixes in v0.9.3 protect a downloader that can run outside browser-mediated controls. The Docker changes in v0.9.0 reduce what network callers can control and require authentication by default. The v0.9.4 security overview reports additional SSRF and configuration-gate fixes.
| Version | Area | Documented change | Operational meaning |
|---|---|---|---|
| 0.9.4 (23 September 2026) | Security overview | Fixes for two SSRF paths and an untrusted-configuration bypass; robots.txt and link-preview fetching use the pinning egress proxy, and nested typed objects are rechecked against the untrusted-configuration gate. | Recheck the release notes and deployed configuration because this status is date-sensitive. |
| 0.9.3 | PDF crawling | Per-hop redirect and peer validation, 100 MiB and 2,000-page limits, safer image-output handling, escaped PDF text, and automatic strategy routing. | Large, redirected, malformed, or untrusted PDFs are constrained at their own download boundary. |
| 0.9.0 | Self-hosted Docker API | Authentication enabled by default, loopback binding unless a token is configured, untrusted request bodies, and authenticated artifact retrieval with TTL and storage quota. | Existing HTTP-server deployments may need migration; the in-process Python library was not changed by this release. |
How the PDF security boundary works
Why browser controls were not enough
A Docker API request could select PDFContentScrapingStrategy. The server then used PDFCrawlerStrategy, which fetched the document through Python’s requests path rather than through the browser’s egress and resource controls. A policy applied to browser navigation therefore did not automatically constrain the PDF fetch. The fix is not one global switch; it is a set of checks at the PDF downloader’s trust boundary.
Recommended Free Tools
#1 Best Overall
Redirect validation prevents destination changes
Checking only the URL supplied by a caller is insufficient. A permitted public URL can redirect to an internal address on the next hop. Crawl4AI’s documented PDF handling validates redirect destinations manually, allows at most five hops, and validates the peer IP of the response. Treat the five-hop value as the project’s stated limit, not a universal safe number for every deployment.
Download and page limits contain resource exhaustion
Version 0.9.3 sets max_pdf_bytes to 100 MiB and max_pdf_pages to 2,000. Untrusted Docker request bodies cannot raise those values above the caps. The Docker configuration also changes limits.wall_clock_s to a 300-second default. These controls address oversized files, pathological page counts, and work that never completes; tune surrounding infrastructure as well rather than assuming the defaults fit every workload.
Untrusted bodies cannot choose arbitrary local output paths
The release filters save_images_locally and image_save_dir from untrusted request bodies and forces extract_images off for those bodies. A caller can no longer use a scrape request to select an arbitrary server filesystem destination for extracted images. If your application needs image extraction, make that an explicit, trusted server-side policy with a fixed, non-sensitive directory.
Parsed text is escaped before HTML insertion
Paragraph text extracted from a PDF is escaped before it is placed in cleaned_html. The release also removes a Playground viewer round-trip that could interpret crawled content as live HTML. Escaping protects the presentation and downstream HTML consumers; it does not make an untrusted PDF harmless to the parser itself.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsStrategy selection is wired automatically
In the v0.9.3 Docker server, selecting PDFContentScrapingStrategy routes the request to PDFCrawlerStrategy automatically. You do not need to pair those strategies manually for the documented server path. Confirm the behavior against the exact image or package version you deploy.
Deploying the Docker API securely
Start with the new default posture
The v0.9.0 self-hosted Docker API enables authentication by default and binds to loopback unless a token is configured. Network request bodies are treated as untrusted input. This is a breaking change for the self-hosted HTTP server, not for Crawl4AI’s core in-process Python use.
- Identify the interface you run. If your application imports the Python library in the same process, the v0.9.0 HTTP-server changes do not automatically apply. If another container, host, or network client calls the Docker API, apply the server migration steps.
- Keep loopback binding unless remote access is required. For a remote caller, configure a token deliberately and place the service behind your normal network controls rather than exposing an unauthenticated listener.
- Assume every request body is hostile. Do not let client-supplied configuration decide filesystem destinations, extraction behavior, limits, or other server internals.
- Handle artifacts as capabilities. Screenshot and PDF output is moved to artifact identifiers fetched through an authenticated endpoint, with a time-to-live and storage quota. Protect the artifact endpoint and monitor expiry and quota behavior.
- Review the migration guidance for your exact version. Configuration names, container images, and endpoint details can change; do not copy a v0.9.0 assumption into a later deployment without checking the project’s current release documentation.
What the release does not guarantee
The release notes describe project changes, not an independent penetration test, performance benchmark, or compliance certification. AWS’s general CloudFront guidance puts the boundary clearly: “Security is a shared responsibility between AWS and you.” The same principle applies when you place Crawl4AI behind a cloud load balancer or CDN: the provider protects its infrastructure, while you still own identity, network exposure, secrets, data handling, and runtime configuration.
Limits and controls to configure around
| Risk | Crawl4AI control | What you should add operationally |
|---|---|---|
| Redirect to an internal or forbidden host | Manual destination checks, peer-IP validation, maximum five PDF redirect hops | Restrict outbound routes and log every redirect decision. |
| Huge or page-heavy PDF | 100 MiB and 2,000-page caps | Set queue limits, reject early where possible, and budget memory per worker. |
| Never-ending processing | limits.wall_clock_s default of 300 seconds in the Docker configuration |
Use worker timeouts, cancellation, and a bounded retry policy. |
| Arbitrary server writes | Untrusted bodies cannot set save_images_locally or image_save_dir; extraction is forced off |
Use fixed writable volumes and least-privilege filesystem permissions. |
| Markup injection from extracted text | PDF paragraph text escaped before cleaned_html |
Keep escaping and output-context encoding in downstream templates and indexes. |
| Expired or excessive output artifacts | Authenticated artifact endpoint with TTL and storage quota | Copy required artifacts to controlled storage and alert on quota pressure. |
Processing untrusted PDFs at scale
Isolate the parser
Apache PDFBox states, “Processing untrusted PDFs is supported, but only to a defined extent.” Its security guidance identifies remote code execution, privilege escalation or sandbox escape, unauthorized data access, and malformed-document exhaustion of CPU, memory, recursion, or processing time. Use a separate worker or container, a non-privileged account, strict outbound policy, memory and CPU limits, and a killable timeout. Do not treat a size cap as a substitute for sandboxing.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Rank #3
Harden the PDF application environment
ASD hardening guidance warns that default or unapproved PDF-application settings can create an insecure environment. Follow vendor and ASD hardening guidance, prevent users from weakening those settings, and block PDF applications from creating child processes where your workload permits. Apply the most restrictive applicable guidance when recommendations conflict.
Keep software-supply-chain controls in scope
The UK Software Security Code of Practice is a voluntary baseline of 14 principles for software supplied to business customers. It is guidance, not evidence that Crawl4AI or a deployment is certified. Use it to structure requirements for dependency updates, vulnerability response, build provenance, access control, and recovery.
Separate scraping permission from parser safety
These controls do not decide whether you may crawl a particular site or document. Check the target’s terms, robots policy, contracts, copyright obligations, privacy requirements, and applicable law separately. A technically hardened crawler can still be used against a target you are not authorized to access.
Local processing versus hosted PDF services
Self-hosting keeps documents inside infrastructure you control, but makes you responsible for patching, isolation, egress policy, storage, logs, and incident response. A hosted service shifts some operations to a provider without removing data-governance decisions.
Adobe documents that its server-side PDF Services and PDF Embed components run in Adobe Document Cloud on AWS infrastructure in US-East and EMEA, with a customer-selectable processing region. It also documents temporary caching of user-generated content, TLS 1.2 or greater for content in transit, and permission settings that can prevent processing. Password-protected PDFs cannot be processed unless the password is known and the author has authorized removal of protection. Confirm region, retention, access controls, and document permissions before uploading sensitive files.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| The Docker endpoint is unreachable from another host. | Loopback binding is the v0.9.0 default. | Keep loopback for local use, or configure remote binding deliberately with token authentication and network restrictions. |
| A request is rejected after selecting PDF scraping. | The file exceeds 100 MiB, exceeds 2,000 pages, violates redirect or peer checks, or runs past the wall-clock limit. | Inspect the server decision and source document; do not raise limits from an untrusted request body. Split or pre-approve documents in a trusted workflow. |
| Images are not extracted to the requested directory. | Untrusted bodies have extract_images forced off and cannot set local-write fields. |
Use a trusted, fixed server policy and a controlled volume. |
| Extracted text appears as literal markup. | PDF text is escaped before insertion into cleaned_html. |
Render it as text or apply context-appropriate, safe formatting in your own presentation layer. |
| An artifact link has expired or returns no output. | The artifact TTL elapsed or the storage quota was exceeded. | Fetch artifacts promptly through the authenticated endpoint and monitor quota; persist important results in your own controlled store. |
| A redirect that worked in a browser is blocked by the scraper. | The destination or connected peer fails the PDF path’s SSRF checks. | Verify the redirect chain and destination ownership. Do not bypass the check merely to make one URL work. |
| The v0.9.4 security behavior is absent. | The running image or package is older than the release you reviewed. | Check the actual deployed version, not just the client dependency, and re-read the matching migration and security notes. |
Performance, reliability, and cost decisions
The documented material contains no independent throughput, latency, or incident-rate benchmark. The explicit 100 MiB, 2,000-page, five-hop, and 300-second values are software controls, not measurements of capacity. Expect larger PDFs and multi-hop redirects to consume more worker time and memory; size queues and concurrency around your own documents, then measure with representative workloads.
- Use a bounded queue so one page-heavy document cannot occupy every worker.
- Record verdicts, redirect destinations, byte and page counts, elapsed time, and artifact-expiry outcomes without logging document secrets.
- Retry transient network failures selectively; do not blindly retry a deterministic policy rejection or a blocked internal destination.
- Keep parser workers disposable so a crash or memory spike does not compromise the API process.
Or skip the browser setup: ScreenshotNeo
If your goal is a clean website screenshot or PDF rather than operating a browser and PDF-ingestion stack, ScreenshotNeo provides a single website-screenshot API and an MCP server. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the result with X-Page-Verdict and X-Billed headers.
Use the ScreenshotNeo API documentation for the complete option list and authentication details. The following requests are runnable; replace the key and target URL.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Options relevant to scraper and document workflows
- Full-page capture with lazy images loaded; capture one element by CSS selector; dark mode; 12 device presets or any viewport; and retina scale.
- PDF output with paper size, margins, landscape mode, and page ranges.
- HTML/CSS to image; custom CSS and JavaScript; click an element before capture; hide selectors; and wait for a selector, delay, or network idle.
- Block ads, trackers, requests, or resource types; set custom headers, cookies, user agent, and Authorization; and set timezone and geolocation.
- Transparent backgrounds and image resizing.
- Choose a cache TTL, create signed links for public
<img>tags, submit asynchronous jobs with signed webhooks, capture up to 100 URLs per bulk call, query usage, and use the OpenAPI specification. - Parameter names used by other screenshot APIs also work, which can simplify migration.
- The MCP server exposes
take_screenshot,get_page_info, andcapture_pdffor Claude, Cursor, and other MCP clients.
Plans include Free (1,000 shots per month, no card), Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000). Yearly billing gives two months free, and every feature is available on every plan.
Best Value
When you want the browser setup handled for you, create a free ScreenshotNeo account with 1,000 screenshots a month and no card required.
Release-check checklist
- Confirm the deployed Crawl4AI version; v0.9.4 was listed as latest on 23 September 2026.
- Verify authentication, binding, token handling, and artifact TTL/quota on the Docker server.
- Test redirects to public, private, and unexpectedly redirected destinations.
- Exercise byte, page, and wall-clock limits with harmless fixtures.
- Confirm untrusted bodies cannot set local output paths or re-enable image extraction.
- Run malformed and password-protected document cases inside an isolated worker.
- Review downstream HTML rendering so escaped PDF text is never treated as executable markup.
- Document data region, retention, and permissions before using any hosted PDF service.
Conclusion
Crawl4AI’s recent work is best understood as defense in depth. The PDF downloader now validates each redirect path, constrains bytes, pages, and time, blocks untrusted filesystem choices, and escapes extracted text. The Docker API defaults to authentication and restricted binding, while artifacts move behind an authenticated, expiring endpoint. Those controls reduce exposure; isolation, least privilege, outbound filtering, patching, and authorization to crawl remain your responsibility.
Frequently Asked Questions
Does the v0.9.0 change affect Crawl4AI’s in-process Python library?
The release notes describe the breaking changes for the self-hosted Docker HTTP server; they state that the core in-process Python library was unchanged.
Are Crawl4AI’s v0.9.4 security fixes independently audited?
The available material reports project release and security-overview changes, not an independent audit or benchmark.
Does a PDF size limit make processing safe by itself?
No. Size and page caps reduce one exhaustion vector, but parser isolation, memory and CPU limits, timeouts, outbound controls, and hardened PDF settings are still needed for untrusted documents.
Can I assume a hosted PDF service keeps documents in my preferred jurisdiction?
No. Confirm the provider’s processing region, temporary caching, retention, access controls, and document-permission behavior before uploading sensitive files.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




