Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
For a news website that permits automated access, begin with its API, RSS or Atom feed, JSON feed, or sitemap. If none meets your needs, use Python Requests to fetch one allowed page and Beautiful Soup to extract the fields you need. Check the publisher’s access rules and reuse terms first: robots.txt is an important crawler-access signal, but it does not grant copyright or other reuse rights.
This guide builds a cautious scraper for article links and headlines, then explains how to adapt it for article metadata, validate and store records, and troubleshoot common failures.
Before scraping: define the task and check permission
Decide what you need
Choose the publisher and section, the specific pages to fetch, and the fields to retain. A small, explicit schema is easier to validate and less likely to collect information you do not need. For news articles, typical fields are the canonical URL, headline, publication or update time, byline, section, summary, article body, and retrieval time.
Start with one permitted page. Inspect its public HTML and the publisher’s current rules before expanding to pagination or other sections. Do not scrape merely because a page is visible in a browser.
#1 Best Overall
Check robots.txt and the publisher’s terms
Python’s urllib.robotparser reads robots rules and can answer whether a user agent may fetch a URL. It can also expose crawl-delay, request-rate, and sitemap information when the file provides them. A robots rule is about crawler access and traffic; it does not settle copyright, privacy, licensing, database rights, or terms-of-service questions. Read the publisher’s relevant policies as well.
The example below uses https://example-news-site.test/robots.txt only as an illustrative target. Replace the example domain and contact details with your own, and stop if the target publisher prohibits the requested access. Do not bypass a paywall, login, CAPTCHA, access control, or an explicit prohibition.
Prefer an API, feed, or sitemap when available
Before parsing article HTML, look for an official publisher API, RSS or Atom feed, JSON feed, or sitemap. Structured publisher data is generally less fragile than page selectors and may define authentication, quotas, and permitted reuse. Follow the provider’s documented limits and reuse terms.
Use Requests and Beautiful Soup when the permitted content is present in ordinary HTML. If you need crawl orchestration, link-following, and pagination, Scrapy may be a better fit. Use browser automation only when the content genuinely requires client-side rendering and the publisher’s rules permit that access; it is heavier and more complex than an HTTP request.
Rank #2
A conservative Python scraper for news listing pages
Install the two dependencies with python -m pip install requests beautifulsoup4. This runnable example checks robots.txt, identifies the client, uses a finite timeout, checks the HTTP response, extracts links and headlines from article cards, and records when it retrieved the page.
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
URL = "https://example-news-site.test/news"
USER_AGENT = "ExampleResearchBot/1.0 (+mailto:[email protected])"
TIMEOUT_SECONDS = 15
parsed = urlparse(URL)
robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
robots = RobotFileParser(robots_url)
robots.read()
if not robots.can_fetch(USER_AGENT, URL):
raise RuntimeError(f"robots.txt does not allow {USER_AGENT} to fetch {URL}")
response = requests.get(
URL,
headers={"User-Agent": USER_AGENT},
timeout=TIMEOUT_SECONDS,
)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
retrieved_at = datetime.now(timezone.utc).isoformat()
articles = []
for card in soup.select("article"):
link = card.select_one("a[href]")
headline = card.select_one("h1, h2, h3")
if not link or not headline:
continue
articles.append({
"url": urljoin(URL, link["href"]),
"headline": headline.get_text(" ", strip=True),
"retrieved_at": retrieved_at,
})
for article in articles:
print(article)
The article and heading selectors are examples, not a universal news-site template. Inspect the target’s permitted HTML and change the selectors to match its actual structure. Keep the robots check, timeout, status check, whitespace normalization, and retrieval timestamp even when you change the parsing logic.
What the example does not do
It fetches one listing page and extracts cards. It does not follow article links, implement pagination, infer publication dates, or determine whether reuse of the extracted material is allowed. Add those behaviors only when needed and permitted.
Free tools Windows power users keep installed
One-click scans. No signup required.
Extracting fields from an article page
Once you have an allowed article URL, fetch it with the same identifying user agent, timeout, and status check. Prefer semantic HTML and publisher-supplied structured data such as JSON-LD when present. These examples illustrate common patterns; inspect each target page rather than assuming a selector or JSON-LD shape.
article_url = articles[0]["url"]
article_response = requests.get(
article_url,
headers={"User-Agent": USER_AGENT},
timeout=TIMEOUT_SECONDS,
)
article_response.raise_for_status()
article_soup = BeautifulSoup(article_response.text, "html.parser")
headline_tag = article_soup.select_one("h1")
canonical_tag = article_soup.select_one('link[rel="canonical"]')
record = {
"url": urljoin(article_url, canonical_tag["href"])
if canonical_tag and canonical_tag.get("href") else article_url,
"headline": headline_tag.get_text(" ", strip=True) if headline_tag else None,
"byline": (article_soup.select_one('[rel="author"]') or {}).get_text(" ", strip=True)
if article_soup.select_one('[rel="author"]') else None,
"retrieved_at": datetime.now(timezone.utc).isoformat(),
}
print(record)
The byline selector above is only a possible semantic convention; sites also use other markup or JSON-LD. Add extraction for publication time, section, summary, and body only after checking the page structure and the publisher’s reuse terms. Avoid taking all page text indiscriminately: navigation, related stories, advertisements, and reader comments may be mixed in with article content.
Make the crawl polite and bounded
- Use an informative User-Agent with a real contact address, and set a finite timeout on every request.
- Fetch one page at a time where possible. Keep concurrency conservative, and honor any crawl-delay, request-rate, API quota, or other publisher instruction that applies.
- Check status codes before parsing. On transient failures, use a limited retry policy with backoff; stop rather than retrying indefinitely or hammering a failing site.
- Cache fetched pages and avoid requesting the same URL repeatedly. Set a clear maximum page count and pagination bound.
- Stop on repeated errors, an explicit prohibition, or a change in access conditions. Do not try to evade blocks.
Reading crawl hints with RobotFileParser
After reading the file, RobotFileParser provides methods such as crawl_delay(user_agent), request_rate(user_agent), and site_maps() for directives present in the robots file. Treat these as crawl hints to inspect and honor where applicable; the absence of a value is not permission to crawl aggressively.
Validate, deduplicate, and save records
Before a record enters a dataset, reject or quarantine it if it lacks a canonical URL or headline. Normalize whitespace and timestamps consistently, and deduplicate articles by canonical URL so that query strings or listing-page variations do not create unnecessary duplicates. Keep the source URL and the time you retrieved each record.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →For a durable pipeline, store the publisher, byline, publication time, retrieval time, parser version, and any license metadata alongside the fields you collect. Log each run, including the pages attempted, records accepted or rejected, and errors. If selectors stop matching after a template change, the run log and validation checks help catch missing data instead of silently saving empty records.
Choosing the right tool
| Approach | Use it when | Trade-off |
|---|---|---|
| Publisher API or feed | The publisher offers structured data that covers your fields and permits your intended use. | Follow its authentication, quotas, schema, and reuse terms. |
| Requests and Beautiful Soup | The permitted content is available in static HTML and the task is a small or controlled extraction. | Selectors depend on page markup and need maintenance when templates change. |
| Scrapy | You need crawl orchestration, pagination, and a larger managed crawl. | It adds framework setup; it does not remove the need to respect access and reuse rules. |
| Browser automation | Permitted content is rendered client-side and a plain HTTP fetch cannot access it. | It uses a browser and is more resource-intensive; it is not a way to bypass controls. |
When a screenshot is enough instead of scraped text
If the goal is to archive or inspect how a permitted page looks, rather than extract article text into structured records, a screenshot may be more appropriate than an HTML scraper. ScreenshotNeo is a website screenshot API and MCP server for developers. Its captures can return PNG, JPEG, WebP, or PDF; options include full-page capture, CSS-selector element capture, viewport and device settings, and waiting for a selector or network idle. A screenshot is not a substitute for permission to access a page or reuse its contents.
Or skip the browser setup
For a screenshot rather than scraped fields, make one GET request. This cURL example saves a WebP capture of the target page:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example-news-site.test/news -o shot.webp
See the ScreenshotNeo documentation for API parameters and response details. Cookie banners are accepted before capture and more than 60 known consent platforms, newsletter popups, and chat widgets are removed; each step can be turned off. Bot checks, blank pages, and failed loads are not billed. An MCP server lets AI agents use screenshot tools, and the free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common failures
Robots check says the URL is disallowed
Confirm that you constructed the robots URL from the target’s scheme and host, that the checked URL is the one you intend to fetch, and that the user agent matches the one you send in Requests. If the publisher’s rules disallow access, stop; do not change user agents to evade the rule.
Best Value
The request times out or returns an HTTP error
A timeout prevents the client from waiting indefinitely; it does not make a failing page safe to retry without limit. Check the URL and response status, slow down, and use a small bounded retry with backoff only for transient errors. Stop on repeated failures or a publisher-side block.
The scraper returns no cards or empty headlines
The page may use different HTML than the illustrative selectors, or its content may be rendered client-side. Inspect the permitted HTML, update configurable selectors, and validate the result before saving. If structured data or a feed exists, prefer that. Consider browser automation only if the publisher’s rules allow it.
Recommended Free Tools
Headlines or dates are duplicated or inconsistent
Use the canonical URL where available and deduplicate on it. Normalize whitespace and parse timestamps according to the publisher’s supplied format, retaining the original source value if normalization is uncertain. Keep retrieval time distinct from publication or update time.
The site changes its template
Recheck selectors and permissions when records unexpectedly lose fields or page structure changes. Quarantine invalid records and review the run log instead of treating a successful HTTP response as proof that extraction succeeded.
Frequently asked questions
Does robots.txt mean I have permission to republish article text?
No. It addresses crawler access and traffic, not copyright, licensing, privacy, database rights, or terms of service. Determine reuse rights separately and obtain permission or a license where needed.
Can I use this method for every news website?
No. The code is a starting point for a permitted page whose structure matches the selectors you configure. Publishers set different access rules and page templates, and those can change.
What should I do if an article is behind a paywall or CAPTCHA?
Do not bypass it. Use a publisher-authorized API or access route, or do not collect that content.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

