Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To collect data from a website, first define the pages and fields you need, then use the simplest supported access route: an official API or feed when available, otherwise fetch pages and extract their HTML. Validate and store the results in a format your next tool can use. If the data appears only after JavaScript runs, inspect the browser’s network requests before reaching for browser automation.
Plan the collection before writing code
A useful collection starts with a narrow specification, not a crawler. Write down what you are collecting and why, then decide how the result will be checked and used.
- Scope: Identify the site, page types, and relevant paths. Avoid following every link on a domain when only a catalog, directory, or set of articles is needed.
- Fields: Name each required field, such as title, author, price, date, or canonical URL. Decide how to represent missing values and inconsistent formats.
- Frequency: Determine whether this is a one-time collection or a recurring job. The necessary controls and storage differ if records are refreshed regularly.
- Output: Choose JSON Lines, CSV, XML, or a database according to the downstream application. JSON Lines is convenient for one-record-per-line processing; CSV is broadly usable in spreadsheets.
- Permission and access: Check the site’s terms, supported interfaces, and applicable access rules before collecting.
Keeping the job limited to the fields and pages needed makes it easier to validate, maintain, and run responsibly.
Choose the right way to access the data
| Approach | Use it when | Trade-off |
|---|---|---|
| Official API or feed | The site offers a documented interface that includes the fields you need. | Available fields, access conditions, quotas, and update cadence depend on the site. |
| HTTP client and HTML parser | The required content is present in the page’s initial HTML response and the collection is relatively small. | You may need to implement pagination, retries, scheduling, and export yourself. |
| Scrapy | You need a repeatable crawl with selectors, pagination, request controls, and structured exports. | It introduces framework structure, but combines crawling and extraction features. See the Scrapy documentation. |
| Headless browser | The data genuinely depends on browser execution, or you need a browser-rendered result. | It requires browser setup and more machinery. First check whether the data comes from a request you can reproduce directly. |
| Hosted extraction API | Managed execution and dataset export suit the job. | Compare provider coverage, data quality, terms, cost, and program availability; product documentation alone does not establish comparative performance. |
Scrapy’s documentation covers crawling through APIs as well as pages. For HTML parsing, it describes Scrapy selectors, Beautiful Soup, and lxml; Beautiful Soup is a parser suited to malformed markup, while lxml handles HTML and XML. A parser extracts data from a response; it does not by itself provide a full crawl and scheduling workflow. See the Scrapy selector documentation.
#1 Best Overall
Use an API or feed when one fits
Search the site’s own documentation for an API, export, or feed before parsing its page markup. Confirm that the interface supports your fields and intended frequency, and follow its authentication, quota, and terms. A supported interface can return structured data directly and reduce dependence on page layout, but its capabilities are site-specific.
Do not infer permission from the mere existence of an endpoint. Use documented access routes and assess the site’s terms and applicable requirements separately.
Fetch and parse ordinary HTML
When content is already in the initial HTML response, a lightweight HTTP client can fetch it and a parser can select the relevant elements using CSS or XPath. Inspect the actual response before choosing selectors: a page’s visual appearance does not guarantee that its content is present in the downloaded HTML.
Use selectors tied to the content’s meaning or stable structure where possible. Extract only the fields defined in your plan. Keep the source URL with each record so you can trace where it came from and diagnose changes.
For a small one-off job, a client and parser may be enough. If the job needs pagination, controlled request scheduling, exports, and repeated runs, a crawling framework such as Scrapy may keep those pieces together. Its documentation shows spiders that select fields, follow a next-page link, and export JSON Lines.
Follow pagination without crawling the whole site
Build traversal around the pages that actually contain records. A typical crawler requests a starting page, extracts its records, finds the relevant “next” link, and repeats until there is no next page or the defined boundary is reached.
- Choose an explicit starting URL or a known set of listing pages.
- Extract the records and the next-page link using selectors verified against the response.
- Follow only that relevant link pattern; do not recursively visit unrelated site navigation or outbound links.
- Stop at a known end condition, such as no next link, a chosen date boundary, or a specified page set.
Scrapy’s tutorial demonstrates extracting selected fields, following pagination, and yielding items. Its feed exports support JSON, CSV, and XML; item pipelines can be used for storage and processing. See the Scrapy tutorial and feed export documentation.
Diagnose content loaded by JavaScript
If a field is visible in a browser but absent from the initial HTML, treat that as a source-discovery problem first. Open the browser’s developer tools, inspect the Network panel, reload the page, and look for the request whose response contains the missing data. It may be a JSON or text endpoint, or data embedded in JavaScript.
- Compare the browser-visible content with the initial page response.
- In the Network panel, identify requests made as the missing content appears.
- Inspect the relevant response and determine whether it contains the needed fields.
- If practical and permitted, reproduce that request and parse its response in your collection code.
- Use a headless browser when reproducing the request is impractical or when browser-rendered output is itself required.
Scrapy’s documentation on dynamic content recommends finding the data source and extracting from it. It also shows a Playwright integration example for cases that need browser rendering. See Scrapy’s dynamic-content guidance.
Browser automation is not automatically the best solution just because a page uses JavaScript. If the underlying request is stable and allowed for your use, parsing its response can avoid browser setup. Conversely, browser execution may be the appropriate route when the page’s rendered state matters or the request cannot reasonably be reproduced.
Rank #3
Validate and store the records
Extraction is not complete when a selector returns text. Before using the result, normalize and inspect it.
Recommended Free Tools
- Normalize values: Use consistent field names, whitespace, date formats, and URL forms.
- Handle missing data: Decide whether to retain a record with a missing field, mark it, or reject it. Do not silently substitute a value.
- Check record shape: Confirm required fields exist and have plausible types or formats.
- Keep provenance: Store the source URL and, where useful, collection time so records can be traced and changes investigated.
- Choose storage to fit the job: Use JSON Lines, CSV, or XML for file-based workflows, or an item pipeline or other destination for ongoing processing.
There is no universally best database or retention period for every collection. Choose based on volume, update patterns, and how the data will be used downstream.
Keep collection controlled and responsible
Review the target website’s terms and documented access routes. Check its robots.txt instructions and configure your crawler to honor applicable rules. Keep request rates proportionate, avoid collecting unnecessary fields, and do not attempt to defeat access controls.
Robots.txt is not a security barrier or a complete legal decision. Google describes it as a way to guide crawler access and manage crawler traffic; it does not enforce behavior or reliably hide a page from search results. Google notes that a blocked URL may still be indexed if linked elsewhere, and points to password protection or noindex when the goal is preventing search appearance. Those are Google Search behaviors, not permission to collect the page’s data. See Google’s robots.txt guidance.
Whether a particular collection is permitted can depend on the fields, website, jurisdiction, and intended use. The 2024 paper on web scraping for U.S.-based social science research discusses legal, ethical, institutional, and scientific considerations rather than establishing a rule for every country or project. See the paper’s publication page.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Or skip the browser setup
If your task is to capture a page as an image or PDF rather than build a structured dataset, ScreenshotNeo offers a website screenshot API and MCP server. It is not a replacement for extracting fields into records; it is useful when a clean rendered capture is the output you need. One GET request returns a screenshot or PDF. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.
Sign up for 1,000 free screenshots a month, with no card required.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common collection problems
The page looks right in the browser, but fields are missing
The content may be populated after the initial response. Inspect the Network panel and find the response carrying it; parse that response if practical, or use a headless browser when browser execution is necessary.
Free tools Windows power users keep installed
One-click scans. No signup required.
A selector returns no results
Check the downloaded response rather than relying on the visual page. The selector may target browser-generated content, or the markup may have changed. Verify the element’s structure and update the selector to match the current response.
Pagination stops early or repeats pages
Inspect the extracted next-page URL on each iteration. Make sure it points to the next relevant listing page, and define a stop condition. Avoid broad link-following rules that revisit the same page or wander into unrelated sections.
Best Value
Records contain inconsistent or malformed values
Normalize values after extraction and validate required fields before storing records. Preserve the source URL so you can check whether the issue is in the page, selector, or normalization step.
The crawl sends more requests than expected
Constrain the allowed pages and pagination path. Scrapy documents download delays, per-domain concurrency limits, and AutoThrottle for controlling request behavior; consult its current AutoThrottle documentation and settings reference for configuration details.
Performance, reliability, and cost considerations
For a small collection, a direct HTTP request and parser can minimize moving parts. A crawler framework adds structure for pagination, request controls, and exports, which can help make recurring work manageable. A headless browser is justified when the task requires browser execution, but first check for a usable underlying response. These approaches have different setup and maintenance costs; the right choice depends on the data source and the repeatability required.
Reliability comes from bounded scope, controlled request rates, clear stop conditions, validation, and enough source context to investigate failures. Set expectations for changed markup or unavailable pages, and design the job to identify incomplete or malformed output rather than treating every run as successful. Specific legal permission, quotas, and service costs vary by site or vendor; confirm them for the collection you plan to run.
Frequently Asked Questions
Does robots.txt give permission to collect a website’s data?
No. It communicates crawler preferences; terms, access controls, applicable law, and the intended use need separate consideration.
Should I use a headless browser for every JavaScript website?
No. First inspect network requests for the response that supplies the data. Browser automation is for cases where reproducing the request is impractical or rendered output is required.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsWhich output format should I choose?
Choose for the next step in your workflow: JSON Lines, CSV, and XML are supported Scrapy feed formats, while a database or pipeline can suit ongoing storage.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

