Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use the least invasive, most authorized route first. Check for an official API, RSS feed, sitemap, or permission process; read the site’s terms and privacy policy; inspect the correct robots.txt; then fetch only the article pages you need at a modest rate. Parse the returned HTML with a normal HTTP client and BeautifulSoup for a few known URLs, or use Scrapy for a bounded crawl. Keep extraction separate from storing, analyzing, or republishing the text, because those activities can have different legal and policy consequences.

What “scraping articles” means

Scraping usually means collecting data from a page you already know, while crawling means discovering and following links to find more pages. A project can do both: use a bounded crawler to discover article URLs, then scrape selected fields such as the title, author, publication date and body. Define that boundary before writing code so a one-page task does not become an uncontrolled crawl.

Plan the collection before you send a request

Write a narrow scope

  • Record the target host and the URL pattern that identifies an article.
  • List the fields you actually need: for example, headline, byline, date, canonical URL and body paragraphs.
  • State the purpose, retention period, storage location and who can access the results.
  • Set a maximum number of pages and a stop condition for errors or unwanted responses.

A small sample is easier to validate and less burdensome to the publisher than a recursive crawl. If the task is research or commercial monitoring, document the jurisdiction and intended downstream use before collection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Look for an authorized structured source

Search the publisher’s documentation and footer for an API, RSS or Atom feed, sitemap, downloadable dataset, or a contact address for permission. The Carpentries’ Web Scraping with Python: Hello-Scraping lesson recommends checking whether structured access exists and asking the organization when a special agreement may be appropriate. An API or feed is often more stable and less expensive for both sides than repeatedly parsing page layouts.

Check terms, privacy and robots.txt independently

Read the site’s policies

Review the terms of service, privacy policy and any developer or data-licensing page. Public visibility does not by itself answer whether automated collection, storage, redistribution or commercial use is allowed. Copyright, privacy obligations, access controls and jurisdiction can apply to different stages of your workflow. For substantial research or business use, obtain permission or advice from a qualified legal or institutional source.

Interpret robots.txt correctly

Fetch the file at the target origin’s root, such as https://example.com/robots.txt, and inspect the rule for the user agent you will identify. Google’s specification explains that a robots.txt file applies to the same host, protocol and port that serve it; a file on one subdomain does not automatically govern another. Its rules are a policy signal, not a license or a complete legal answer. The Carpentries lesson puts the practical principle plainly: “To avoid legal or ethical issues, it’s essential to check both the TOS and the site’s robots.txt file before scraping.”

Do not assume that a permissive robots file overrides an express prohibition. Reuters Connect’s Platform Terms and Conditions, last updated September 2024, expressly prohibit scraping and automated collection of platform content without prior written consent and require compliance with exclusionary protocols. Treat every target as its own policy decision.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fetch a few known articles with Python and BeautifulSoup

The following example is for static HTML where the article text is present in the response. It uses a descriptive user agent, a timeout, a delay between requests and a bounded URL list. Replace the selectors after inspecting a few pages from the same site.

import time
from urllib.parse import urlparse

import requests
from bs4 import BeautifulSoup

URLS = [
    "https://example.com/news/first-article",
    "https://example.com/news/second-article",
]
HEADERS = {
    "User-Agent": "ResearchArticleCollector/1.0 (contact: [email protected])"
}


def extract_article(url):
    response = requests.get(url, headers=HEADERS, timeout=30)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")

    title = soup.find("h1")
    author = soup.select_one("[rel='author'], .author, .byline")
    date = soup.select_one("time[datetime], time, .published, .date")
    body = soup.select_one("article, [itemprop='articleBody'], .article-body")

    if body is None:
        raise ValueError(f"No article container found at {url}")

    paragraphs = [p.get_text(" ", strip=True) for p in body.find_all("p")]
    return {
        "url": url,
        "title": title.get_text(" ", strip=True) if title else None,
        "author": author.get_text(" ", strip=True) if author else None,
        "published": date.get("datetime") if date and date.has_attr("datetime") else (
            date.get_text(" ", strip=True) if date else None
        ),
        "text": "nn".join(p for p in paragraphs if p),
    }


for url in URLS:
    if urlparse(url).scheme not in {"http", "https"}:
        continue
    try:
        record = extract_article(url)
        print(record)
    except (requests.RequestException, ValueError) as exc:
        print(f"Skipping {url}: {exc}")
    time.sleep(2)

BeautifulSoup’s find(), find_all(), CSS selectors and text methods are demonstrated in the Carpentries’ instructor lesson. Save structured records rather than unlabelled text so you can trace each value to its source URL and retrieval time.

Validate selectors and records

  • Open several articles, including an older page and one with missing metadata.
  • Confirm that the selected container excludes navigation, “related stories,” comments and newsletter forms.
  • Check that dates retain their timezone or original string when no normalized value is available.
  • Compare the extracted title and first paragraphs with the rendered page.
  • Log HTTP status, final URL, content type and parsing failures without storing unnecessary personal data.

Layouts change. A selector that works today can silently return an empty or contaminated record tomorrow, so fail loudly when a required field is absent.

Use Scrapy for a bounded multi-URL crawl

Scrapy is useful when discovery, scheduling, retries and item pipelines matter. Its downloader middleware includes a robots.txt filter when configured. Scrapy’s documentation states: “This middleware filters out requests forbidden by the robots.txt exclusion standard.”

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
# settings.py
ROBOTSTXT_OBEY = True
USER_AGENT = "ResearchArticleCrawler/1.0 (contact: [email protected])"
DOWNLOAD_DELAY = 2
CONCURRENT_REQUESTS_PER_DOMAIN = 1
AUTOTHROTTLE_ENABLED = True
# spiders/articles.py
import scrapy


class ArticlesSpider(scrapy.Spider):
    name = "articles"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/news/"]

    def parse(self, response):
        for href in response.css("a.article-card::attr(href)").getall():
            yield response.follow(href, callback=self.parse_article)

    def parse_article(self, response):
        body = response.css("article [itemprop='articleBody'] p::text").getall()
        if not body:
            self.logger.warning("No body found: %s", response.url)
            return
        yield {
            "url": response.url,
            "title": response.css("h1::text").get(default="").strip(),
            "author": response.css("[rel='author']::text, .byline::text").get(),
            "published": response.css("time::attr(datetime)").get(),
            "text": "nn".join(x.strip() for x in body if x.strip()),
        }

Keep link discovery bounded with an allowlisted path, a maximum page count and duplicate filtering. Enable ROBOTSTXT_OBEY only as one part of your review; still read the target’s terms and stop if the site indicates that requests are unwanted.

When article text is missing from fetched HTML

If requests returns a shell without the article, first check for an official API, feed or sitemap and ask the publisher about authorized access. The absence of text does not establish that browser automation is required or permitted. A JavaScript-rendered page may also depend on login, a consent choice, an anti-bot check or an access-controlled endpoint. Do not attempt to bypass authentication, CAPTCHAs or technical restrictions. If permission covers browser rendering, use a tool that can run the page in a controlled, low-rate session and continue to honor the same scope and policies.

Protect the site and people represented in the data

  • Identify your client where appropriate and provide a contact address.
  • Request only the pages and fields necessary for the stated purpose.
  • Use delays, low concurrency and off-peak collection; test a small subset first. The U.S. General Services Administration’s July 7, 2021 guidance emphasizes transparency and minimizing impact.
  • Stop when responses show overload, blocking or other signs that collection is causing problems, then seek an authorized route.
  • Minimize personal data, restrict access to stored records and define deletion dates.

Keep extraction separate from reuse

Downloading a page, storing its text, analyzing it and republishing it are separate acts. A research copy may be treated differently from a public database or a commercial feed. Consider whether your output can contain facts, links and metadata rather than expressive article text, and obtain permission for substantial redistribution. The University of Michigan’s copyright guide (May 12, 2022) discusses these distinctions without claiming that all public-web scraping is either legal or illegal.

Choose a method by task

Need Starting point Why
A few known pages with text in returned HTML Python requests plus BeautifulSoup Simple fetching and element-level parsing, as shown in the Carpentries lesson.
Bounded discovery across many article URLs Scrapy with robots middleware enabled Scheduling, item output and robots.txt filtering are built into the framework.
Text absent from fetched HTML Official API, feed or authorized rendered access Structured access is more stable than guessing at browser behavior; permission still governs use.

There is no source-backed benchmark establishing that one of these approaches is universally fastest. Match the tool to page count, rendering needs, policy constraints and maintenance budget.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshooting common failures

403, 429 or repeated timeouts

Cause: rate limits, access policy, an overloaded origin or a blocked client. Fix: stop the run, verify permission and robots rules, reduce concurrency, add a longer delay and contact the publisher. Do not rotate identities or evade a technical block without authorization.

HTTP 200 but no article text

Cause: a JavaScript shell, consent gate, login wall or selector drift. Fix: inspect the raw response, compare it with the rendered page, look for an official feed/API and update selectors only after confirming the page structure and authorization.

Wrong text or duplicated paragraphs

Cause: selecting a broad wrapper that includes navigation, recommendations or mobile duplicates. Fix: narrow the article container, select paragraph nodes, remove known boilerplate and validate against several page templates.

robots.txt cannot be fetched

Cause: DNS, TLS, redirect or server failure. Fix: do not interpret an unavailable file as permission. Pause collection, check the same host/protocol/port manually and ask the site owner for guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrapy follows too many links

Cause: unrestricted pagination, tags or calendar URLs. Fix: constrain allowed_domains, allow only the article path, cap depth or item count, and add explicit URL canonicalization and duplicate checks.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

For a one-call screenshot of a rendered article, ScreenshotNeo accepts a URL and returns PNG, JPEG, WebP or PDF. It removes cookie and consent banners, newsletter popups and chat widgets before capture; bot checks, blank pages, failed loads and timeouts are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

See the ScreenshotNeo API documentation for all options and authorization requirements:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page and element captures, device and retina settings, custom CSS/JavaScript, waits, request blocking, cookies and headers, PDFs, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call and a usage API. Every feature is on every plan. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Further learning

O’Reilly’s Web Scraping with Python, 2nd Edition covers BeautifulSoup, Scrapy and legal and ethical considerations. Use it as a technical reference, while treating each target site’s current terms and permission process as authoritative.

Frequently Asked Questions

Is scraping a public article automatically legal?

No. Public access does not settle copyright, privacy, contract, access-control or jurisdiction questions. Check the site’s terms, robots.txt and intended reuse, and obtain permission when required.

Should I use BeautifulSoup or Scrapy?

Use requests plus BeautifulSoup for a few known static pages. Choose Scrapy when you need bounded discovery, scheduling and item pipelines across many URLs.

Does robots.txt give permission to scrape?

No. It communicates crawler preferences for a specific host, protocol and port. Treat it as one policy signal and review terms and authorization separately.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should I do if the page requires JavaScript?

Look for an official API or feed first. If rendered access is authorized, use a controlled browser workflow; never bypass login, CAPTCHA or other technical restrictions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.