For a small, bounded scrape, run a Python crawler in AWS Lambda; for work that must run longer or continuously, consider ECS or EC2. Before choosing, check the target site’s terms, published API and sitemap, and robots.txt. Identify your crawler, keep requests conservative, and stop when the site denies access. This guide builds a basic, policy-aware Python crawler and explains how to deploy and operate it on AWS.
Choose an AWS compute pattern for the job
There is no universally best AWS service for scraping. Match the runtime and operating model to the workload: a short, independent task can fit Lambda; a sustained or long-running crawl may call for ECS or EC2. AWS Prescriptive Guidance describes Lambda as an option for smaller or modular crawling tasks and EC2 or ECS as possible fits for larger, long-running tasks. Those are architectural choices, not guarantees of speed or lower cost.
| Pattern | When it fits | Trade-offs to consider |
|---|---|---|
| AWS Lambda | Small, bounded jobs, or crawl units that can be split into separate invocations. | Execution duration and resource quotas constrain what fits in one invocation. The AWS Architecture Blog’s June 2020 scraping article states a 15-minute maximum execution time; check current Lambda quotas before relying on that figure. Dependency packaging and startup time matter, especially for browser automation. |
| Amazon ECS | Containerized crawlers needing a longer-lived runtime, custom dependencies, or sustained work. | You take on container deployment and capacity decisions. Choose a launch and scaling approach that matches the job; the appropriate setup depends on workload and operational requirements. |
| Amazon EC2 | A crawler that benefits from a virtual machine with a persistent process or a highly customized runtime. | You manage the instance environment and its lifecycle. Evaluate capacity, patching, scheduling, and monitoring for your use case. |
AWS’s 2021 architecture discussion describes coordinating Lambda tasks with Step Functions for larger serverless crawler patterns. Splitting work can help when each unit has a clear boundary, but introduces orchestration, state, retries, and deduplication to manage. For a browser-rendered target, account for the browser package, memory, startup, and execution time; AWS’s general architecture guidance is not a current, version-specific browser automation recipe.
Decide whether the scraper needs an HTTP endpoint
A scheduled or internally triggered crawler may not need a public HTTP entry point. If a client must invoke a Lambda function over HTTP, Lambda function URLs offer a direct endpoint; API Gateway is the more feature-rich choice when the API needs capabilities such as advanced authentication, throttling, or monitoring. This is an invocation decision, separate from how the crawler accesses websites.
#1 Best Overall
- Durable Carbon Steel: Rack mount screws and cage nuts are made of high-quality carbon steel with a black finish for high strength and dependable durability.
- Easy Installation: Clear metric threads and uniform pitch for better grip. Nylon washers help secure screws and protect equipment surfaces.
- Organized Storage: All parts are packed in a portable storage box for easy organization and access.
- Wide Compatibility: Fits most square-hole racks and cabinets—ideal for server racks, network cabinets, equipment enclosures, and A/V gear.
- 20-Set Kit: Includes 20 mounting screws with nylon washers (M6 x 20 mm) and 20 square cage nuts—40 pieces in total—meeting daily install and replacement needs.
Check permission and crawl rules before coding
Start with the target’s published API, sitemap, terms, and robots.txt. AWS crawler guidance recommends checking robots and sitemap access indications, honoring a stated crawl delay, identifying the crawler with a user agent, and limiting request rate. A missing robots file does not grant blanket permission. A robots file is a crawl instruction, not a legal determination: review the target’s policies and the relevant AWS terms, and assess whether your specific use is permitted in the applicable jurisdiction.
- Prefer an official API or published data feed where one meets your need.
- Request only paths the site permits and only the data you need.
- Use a descriptive user agent with contact information when appropriate.
- Keep request rates low, add timeouts, and use bounded retries with backoff for transient failures.
- Deduplicate URLs so retries or repeated discovery do not create unnecessary traffic.
There is no universally appropriate request rate or retry schedule; follow the target’s published requirements and tune behavior conservatively. If a response is 403 Forbidden, first check that you have permission, the requested path is allowed, and your rate is appropriate. If it remains forbidden, stop crawling that resource. AWS Prescriptive Guidance says: “If none of the above work, you should respect the decision of the website owners and not crawl the page.” Do not try to evade a denial.
Build a small Python crawler for Lambda
This example fetches a single page only after checking the site’s robots rules. It reads links from that page but does not automatically crawl them. Install the dependencies, set the target URL and user agent, then deploy the file and dependencies together as a Lambda package. The example uses a simple HTML parser and standard-library robots parser; it is not a substitute for a target-specific review.
Rank #2
1. Create the function
Save as lambda_function.py. Set TARGET_URL to a page you are authorized to access. The user agent includes a contact placeholder: replace it with a real address before running. The function returns a small JSON result rather than persisting page content.
import json
import os
import time
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
USER_AGENT = os.environ.get(
"CRAWLER_USER_AGENT", "ExampleResearchBot/1.0 (+mailto:[email protected])"
)
TARGET_URL = os.environ["TARGET_URL"]
TIMEOUT_SECONDS = 15
def lambda_handler(event, context):
parsed = urlparse(TARGET_URL)
if parsed.scheme not in ("http", "https") or not parsed.netloc:
raise ValueError("TARGET_URL must be an absolute HTTP or HTTPS URL")
robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
robots = RobotFileParser()
robots.set_url(robots_url)
robots_response = requests.get(
robots_url,
headers={"User-Agent": USER_AGENT},
timeout=TIMEOUT_SECONDS,
)
if robots_response.status_code == 404:
# Missing robots.txt is not permission. Stop for an explicit review.
raise RuntimeError("robots.txt is missing; review site rules before proceeding")
robots_response.raise_for_status()
robots.parse(robots_response.text.splitlines())
if not robots.can_fetch(USER_AGENT, TARGET_URL):
return {"statusCode": 200, "body": json.dumps({"result": "disallowed_by_robots"})}
delay = robots.crawl_delay(USER_AGENT)
if delay is not None:
time.sleep(delay)
response = requests.get(
TARGET_URL,
headers={"User-Agent": USER_AGENT},
timeout=TIMEOUT_SECONDS,
)
if response.status_code == 403:
return {"statusCode": 200, "body": json.dumps({"result": "forbidden_stop"})}
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else None
links = []
for anchor in soup.select("a[href]"):
links.append(urljoin(TARGET_URL, anchor["href"]))
return {
"statusCode": 200,
"body": json.dumps({"url": TARGET_URL, "title": title, "links": links[:50]}),
}
The explicit stop for a missing robots file is a conservative choice in this example, not a universal protocol requirement. Check the site’s access rules and decide whether crawling is allowed before changing that behavior. The code honors a user-agent crawl-delay when one is provided; it does not claim that one request per Lambda invocation is a suitable rate for every target.
2. Package and deploy
- Use a Python version supported by your Lambda runtime. In a clean Linux-compatible build environment, install dependencies into a package directory:
python -m pip install requests beautifulsoup4 -t package. - Copy
lambda_function.pyinto that directory, then create a ZIP from the directory contents. For example:cd package && zip -r ../crawler.zip .. - Create a Lambda function in the AWS Console using a supported Python runtime, upload
crawler.zip, and set the handler tolambda_function.lambda_handler. - Under the function’s environment variables, set
TARGET_URLandCRAWLER_USER_AGENT. Keep secrets out of source code and logs; use an appropriate controlled AWS resource if the crawler needs credentials. - Set a timeout and memory allocation that match the bounded task, then invoke a test event. Review the result and logs before scheduling or exposing the function to callers.
For a larger crawl, do not simply increase parallel invocations without considering the target’s rate limits. Split permitted work into bounded units, persist a deduplication key or crawl state, and coordinate concurrency so the combined request rate remains acceptable.
Rank #3
- Complete Rack Mount Kit: Includes 40 pack M6x16mm cage nuts, screws, and plastic washers, ideal for securing servers in racks or cabinets
- Durable & Corrosion-Resistant: Made of metal with black nickel plating for long-lasting strength and rust prevention, perfect for demanding environments like data centers or industrial setups
- Easy Installation: Spring-loaded cage nuts snap securely into square rack holes, while plastic washers protect equipment surfaces from scratches during tightening
- Universal Compatibility: Designed for standard 19-inch server racks with square mounting holes, ensuring seamless integration with most rack-mountable hardware
- Heavy-Duty Performance: Engineered for durability, these nuts and screws support high-stress applications, from data center servers to industrial AV systems
Operate the crawler without wasting requests
Retries, errors, and durable results
Retry only transient failures, with a limited number of attempts and exponential backoff; do not retry a 403 as though it were a temporary network problem. Treat timeouts separately from successful empty pages. Store only the data needed for the task, protect credentials and extracted data with controls appropriate to their sensitivity, and avoid logging full page bodies or secrets. AWS does not prescribe one security or retention design for every crawler.
Performance and cost
Keep each Lambda unit small enough to finish within its configured runtime and current service quotas. Measure your own latency and memory use for the target and dependencies rather than relying on an assumed throughput. Browser-rendered pages are more resource-intensive than fetching static HTML and may change which compute pattern is practical. AWS charges depend on the current service pricing and your region, networking, storage, configuration, and request volume; estimate costs for the actual design rather than assuming scraping has a fixed per-page price.
Scheduling and invocation
Use an AWS scheduler or an orchestrator to trigger bounded jobs at a cadence the site permits. If a public HTTP call is required, choose a Lambda function URL for a straightforward endpoint or API Gateway when its additional API controls fit the production requirements. Keep the endpoint’s access controls aligned with the sensitivity and intended callers of the scraper.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshoot common failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Lambda reports a timeout | The job, network wait, or dependency startup exceeds its configured runtime. | Reduce scope, use explicit request timeouts, split permitted work into smaller tasks, or evaluate ECS/EC2 for work that must run longer. Check current Lambda quotas before designing around a duration limit. |
| 403 Forbidden | The site denies the path, crawler, or request. | Verify site rules and your rate. If permission is not granted or the response persists, stop; do not evade the denial. |
| Robots check fails | The robots request timed out, returned an error, or the file is missing. | Do not treat a failed check as authorization. Review the site’s access policies and handle the result explicitly before crawling. |
| Import error after deployment | A dependency is absent or packaged at the wrong ZIP path, or the build is incompatible with the runtime environment. | Rebuild dependencies for a compatible environment, confirm the package files are at the ZIP root, and verify the Lambda handler name. |
| Empty or incomplete extraction | The content may be client-rendered, lazy-loaded, or structured differently than expected. | Inspect an authorized page and determine whether an official API is available. If browser rendering is necessary, account for browser packaging and resource needs rather than assuming the static HTML example will capture it. |
| Repeated pages or excess requests | Discovered URLs are duplicated, retries are unbounded, or parallel work is not coordinated. | Normalize and deduplicate URLs, bound retries, and control total concurrency against the target’s rules. |
Or skip the browser setup
If the task is to capture a clean website screenshot rather than extract structured records, ScreenshotNeo provides a screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF; the API documentation is at screenshotneo.com/docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. This is a screenshot service, not a general-purpose structured-data crawler. Sign up for 1,000 free screenshots a month with no card.
Further reading
AWS Prescriptive Guidance’s ESG crawling example illustrates checking robots.txt, allowed paths, crawl delay, and a custom user agent. AWS’s Lambda invocation documentation covers function URLs and API Gateway choices. Review AWS’s legal portal and the target site’s own policies before operating a crawler. For broader Python scraping instruction, Ryan Mitchell’s Web Scraping with Python, 3rd Edition (O’Reilly Media, February 2024) covers parsing, Scrapy, storage, JavaScript and APIs, and legal and ethical considerations; it is not an AWS deployment manual.
Recommended Free Tools
Frequently Asked Questions
Does a robots.txt file decide whether scraping is legal?
No. It communicates crawl rules; legality and contractual permission depend on the target, your use, and applicable jurisdiction.
Best Value
Can I use this example to scrape pages that require a login?
Only if you are authorized and the site permits that access. This example does not implement authenticated sessions.
Is ScreenshotNeo a replacement for a crawler that extracts records?
No. It returns page screenshots or PDFs; use a crawler or an authorized API when you need structured data.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

