October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Beautiful Soup

Data Mining with Web Scraping: Methods and Practical Examples

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Web scraping collects structured records from web pages; data mining cleans and analyzes those records to answer a question. A practical workflow is to choose an appropriate data source, extract a defined set of fields, check and normalize the output, and then analyze it with methods suited to the question. This guide shows the stages with Python and explains how to handle pagination, request pacing, robots.txt, data quality, and common failures.

Scraping and data mining are different stages

Scraping is the collection step: a program fetches pages and extracts fields such as a name, date, category, or price. Data mining comes later. It involves preparing the collected records and applying analysis—such as counts, comparisons, summaries, or text analysis—to find useful patterns.

That distinction matters because extracting a field does not establish that a trend is real, that the data is complete, or that the pages represent a broader population. A useful result depends on the collection scope, the quality of the records, and whether the analysis matches the question.

Choose the right collection method

Before parsing page markup, check whether the publisher provides an appropriate API or dataset. A supported interface can be more stable and easier to use than extracting information from presentation-oriented HTML. Confirm its current documentation, terms, and suitability for your intended use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Approach Best fit Trade-offs
Beautiful Soup or lxml A small, focused extraction from fetched HTML. You control the parsing, but must build any needed fetching, pagination, and storage workflow yourself.
Scrapy Multi-page collection, pagination, structured output, or scheduled crawls. It brings selectors, asynchronous scheduling, exports, pipelines, and request controls, but requires learning more framework concepts.
API or published dataset The site offers a suitable supported data interface. Check the interface’s documentation, terms, fields, and limits; availability and permission are site-specific.

Scrapy documents CSS and XPath selectors and discusses Beautiful Soup and lxml as alternatives. Its walkthrough demonstrates extracting structured records, following pagination, and exporting JSON Lines. See Scrapy at a glance and Scrapy selectors.

Also consider whether the content is rendered dynamically, how often the page structure changes, where output should go, and what access conditions apply. A parser can only extract content present in the response it receives; whether a particular site requires a browser or offers an API must be checked for that site.

Build a small Scrapy spider

Define the records you need before writing selectors. This illustrative spider extracts a name and category from repeated article elements, follows a next-page link, and can be run as a Scrapy project spider. Replace the example URL and selectors with ones that match a source you are allowed to collect from. The example domain and selectors are illustrative, not a tested or authorized target.

import scrapy

class ExampleSpider(scrapy.Spider):
    name = "example"
    start_urls = ["https://example.org/list/1"]

    def parse(self, response):
        for row in response.css("article.record"):
            yield {
                "name": row.css("h2::text").get(),
                "category": row.css(".category::text").get(),
                "source_url": response.url,
            }
        next_page = response.css('a.next::attr("href")').get()
        if next_page:
            yield response.follow(next_page, self.parse)

In a Scrapy project, save the class in the spider module, then run it from the project directory and export records as JSON Lines:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
scrapy crawl example -O records.jl

The output file contains one JSON object per line. Inspect several records before using them: selectors can return null when markup differs, and a selector that matches the wrong element can produce valid-looking but incorrect data. Scrapy’s official tutorial uses the same core pattern—select repeated records, extract fields, follow the next page, and export structured output.

For a one-page parse

Beautiful Soup or lxml can be a lighter choice if HTML has already been fetched and there is no need for crawl scheduling or integrated exports. With either parser, fetching, timeouts, retries, pagination, and storage remain your responsibility. Scrapy’s selector documentation describes how these parsing options compare.

Control crawl pressure and interpret robots.txt

More concurrent requests can increase throughput, but they also increase load on the destination. Scrapy provides download delays, per-domain concurrency limits, and AutoThrottle to control request pace. Set conservative controls appropriate to the site and monitor responses; these controls manage pressure, not permission.

RFC 9309, the IETF’s September 2022 Robots Exclusion Protocol specification, says crawlers are expected to honor parseable robots.txt rules when the file is successfully retrieved. When robots.txt is unreachable because of server or network errors, the RFC says the crawler must assume complete disallow. It distinguishes this from an unavailable response. The specification also states: “These rules are not a form of access authorization.” Read RFC 9309 for the protocol details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots.txt is therefore neither a complete permission grant nor a complete statement of legal rights. Check the site’s terms and applicable rules for your use, jurisdiction, and data. Copyright, privacy, contract, and access questions depend on those particulars; where appropriate, use an official API or licensed dataset instead.

Prepare records before analyzing them

Keep a small, explicit schema. Alongside the extracted fields, retain the source URL and collection date so you can trace a record back to its origin and understand when it was observed. Then perform a validation pass before calculating results.

  • Normalize text consistently, including whitespace and encoding where needed.
  • Convert dates to a consistent representation and units to a consistent scale.
  • Count missing values and decide whether to exclude, retain, or separately classify incomplete records.
  • Identify duplicates, taking care not to collapse distinct records that happen to share a field.
  • Inspect malformed values and outliers against the source page before correcting or removing them.
  • Record which pages and dates were included, and what the collection did not cover.

These steps are practical safeguards, not guarantees that a dataset is complete or representative. Pages can change, records can be repeated across pages, and an extraction may omit content that the selector does not capture. Document those limits when interpreting results.

Match the analysis to the question

Start with the question, then choose a method that answers it without claiming more than the collected records support:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Descriptive questions: count records, summarize numeric fields, or report the distribution of categories.
  • Group comparisons: compare the same measure across defined categories or time periods, while checking that collection coverage is comparable.
  • Text fields: summarize or classify prose only after considering how missing text, duplicates, and changing page content affect the sample.

A count describes the pages and records you collected. It does not, on its own, prove a population-wide trend or explain why a pattern occurred. The collection scope and time window are part of the result, not incidental implementation details.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common scraping problems

Fields are empty or null

The selector may not match the current HTML, the field may be nested differently, or the response may not contain content rendered after the initial page load. Inspect the response and test selectors against a few records. Confirm whether the page needs a different supported data interface or a browser-rendered workflow; do not assume every site behaves alike.

Only the first page is collected

Check that the next-page selector matches a real link and that the link is followed with a relative-URL-aware method such as response.follow. Inspect the final page and confirm that pagination actually continues rather than switching to another mechanism.

Records repeat or look inconsistent

Pages may overlap or the source markup may use different structures for some entries. Preserve source URLs, inspect duplicates before removing them, and validate field formats rather than assuming every extracted value uses the same units or date convention.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Requests fail or the crawl is too aggressive

Check response status and error logs, verify the source is reachable, and reduce concurrency or increase the delay. Scrapy’s AutoThrottle and per-domain concurrency controls can help regulate the request rate. If robots.txt cannot be reached due to server or network error, RFC 9309’s rule is to assume complete disallow; do not treat a technical failure as a reason to push through.

Exported data is not analysis-ready

Open the JSON Lines output and validate its schema, null rates, duplicate rate, and representative values. Fix the extraction or document the treatment of questionable records before drawing conclusions.

Or skip the browser setup

If your task is to capture page images or PDFs rather than build a structured text crawler, ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. For example, here is the cURL call (see the API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up free for 1,000 screenshots a month with no card.

Further reading

Ryan Mitchell’s Web Scraping with Python, 2nd Edition (O’Reilly Media, April 2018) covers Beautiful Soup, crawler construction, Scrapy, storage, cleaning and normalization, language analysis, and legal and ethics topics. Its examples date from 2018, so check current library documentation for version-specific behavior.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.