Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a Google News RSS/XML feed as your input, then parse its <item> elements with Beautiful Soup’s XML parser. The workflow has three separate jobs: your HTTP client downloads the feed, Beautiful Soup builds a parse tree, and your code reads fields such as title, link, and pubDate. Beautiful Soup is a Python library for parsing HTML and XML; it is not a Google News API or a news database.

This guide shows a defensive Python implementation, explains what Google’s Feedfetcher documentation does and does not promise, and covers malformed items, timeouts, namespaces, duplicate stories, and responsible request behavior.

What you need before starting

  • Python 3 and a virtual environment are recommended.
  • The Beautiful Soup 4 distribution, installed as beautifulsoup4.
  • An XML-capable parser. The examples use Beautiful Soup’s xml mode.
  • A Google News RSS/XML URL that you are permitted to request. Feed URL conventions are not documented as a stable public API contract, so treat an endpoint as changeable.

Install Beautiful Soup

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
# .venvScriptsActivate.ps1
python -m pip install beautifulsoup4 requests

Beautiful Soup’s documentation supports the built-in HTML parser and third-party parsers. For RSS/XML, explicitly selecting the XML parser avoids treating the feed as ordinary HTML.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The extraction workflow

  1. Request the feed bytes with a timeout and an honest user agent.
  2. Check the HTTP response before parsing.
  3. Pass the bytes to BeautifulSoup(xml_bytes, "xml").
  4. Find every item element.
  5. Read each child safely, because an item can omit a field.
  6. Normalize and store the records, then apply your own deduplication or date filtering.

Complete Python example

from __future__ import annotations

from datetime import datetime
from email.utils import parsedate_to_datetime
from typing import Any

import requests
from bs4 import BeautifulSoup

FEED_URL = "https://news.google.com/rss?hl=en-US&gl=US&ceid=US:en"


def text_or_empty(item: Any, tag_name: str) -> str:
    """Return a child element's text, or an empty string when it is absent."""
    element = item.find(tag_name)
    return element.get_text(" ", strip=True) if element else ""


def parse_pub_date(value: str) -> datetime | None:
    if not value:
        return None
    try:
        return parsedate_to_datetime(value)
    except (TypeError, ValueError, OverflowError):
        return None


def fetch_and_extract(url: str) -> list[dict[str, Any]]:
    response = requests.get(
        url,
        timeout=(10, 30),
        headers={"User-Agent": "news-feed-reader/1.0"},
    )
    response.raise_for_status()

    soup = BeautifulSoup(response.content, "xml")
    records: list[dict[str, Any]] = []

    for item in soup.find_all("item"):
        title = text_or_empty(item, "title")
        link = text_or_empty(item, "link")
        published_text = text_or_empty(item, "pubDate")
        records.append(
            {
                "title": title,
                "link": link,
                "pubDate": published_text,
                "published_datetime": parse_pub_date(published_text),
            }
        )

    return records


if __name__ == "__main__":
    try:
        stories = fetch_and_extract(FEED_URL)
    except requests.RequestException as exc:
        raise SystemExit(f"Could not download feed: {exc}")

    for story in stories:
        print(story["title"])
        print(story["link"])
        print(story["pubDate"])
        print("-")

The parser reads the feed as bytes, which lets the XML declaration and encoding be honored. The helper checks whether each element exists instead of assuming every response has identical content. The example demonstrates three fields; other feed elements may be present, and their availability can change.

Fetching the feed separately

Keeping network access separate from parsing makes testing easier. Save a response to disk, then pass those bytes to the parser without making another request.

cURL download

curl --fail --location --max-time 30 
  -A "news-feed-reader/1.0" 
  "https://news.google.com/rss?hl=en-US&gl=US&ceid=US:en" 
  -o google-news.xml

Parse a saved file

from pathlib import Path
from bs4 import BeautifulSoup

xml_bytes = Path("google-news.xml").read_bytes()
soup = BeautifulSoup(xml_bytes, "xml")
for item in soup.find_all("item"):
    title_node = item.find("title")
    link_node = item.find("link")
    date_node = item.find("pubDate")
    print({
        "title": title_node.get_text(" ", strip=True) if title_node else "",
        "link": link_node.get_text(" ", strip=True) if link_node else "",
        "pubDate": date_node.get_text(" ", strip=True) if date_node else "",
    })

Optional Node.js retrieval

If another service downloads the feed, it can hand the raw XML to your Python parser or store it for later processing. Node.js 18 or newer includes fetch:

const url = 'https://news.google.com/rss?hl=en-US&gl=US&ceid=US:en';
const response = await fetch(url, {
  headers: { 'user-agent': 'news-feed-reader/1.0' },
  signal: AbortSignal.timeout(30000)
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const xml = await response.text();
console.log(xml);

Choosing and handling feed URLs

Google News examples commonly use regional parameters such as US and India variants. Parameters can affect language, country, and edition. Keep the exact URL in configuration rather than embedding assumptions throughout your code.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google’s official Feedfetcher documentation describes a Google service that retrieves RSS or Atom feeds for Google News and WebSub when users request them through an app or service. It does not establish a supported, stable public Google News feed API for third-party scripts. The documentation also does not promise uptime, pagination behavior, an item limit, or permanent URL conventions.

Do not copy Feedfetcher rules into your scraper

Google says its Feedfetcher ignores robots.txt because it acts directly for a human user and says it should not retrieve most sites’ feeds more than once per hour on average. Those statements describe Google’s own agent. They are not permission for an unrelated program to ignore access rules, and they are not a universal interval for your script. Follow the feed publisher’s terms, use caching, and avoid unnecessary polling.

Make parsing resilient

Missing elements

Use conditional access, as the example does. An absent pubDate should become an empty value or None, not terminate the entire batch.

XML parser unavailable

If Beautiful Soup reports that no XML parser is installed, install an XML-capable parser package supported by your environment and continue to call BeautifulSoup(data, "xml"). Do not silently switch to HTML mode for an XML feed when predictable element handling matters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CDATA and escaped text

Beautiful Soup exposes decoded element text through get_text(). Strip surrounding whitespace, but preserve the resulting title and URL exactly unless your application has a documented normalization rule.

Namespaces and nonstandard elements

RSS extensions can add namespaced elements. Start with the standard item, title, link, and pubDate fields. Inspect a saved response before adding selectors for optional extensions, because names and availability can vary.

Duplicates and changing order

Feeds can repeat a story between requests or reorder items. Use the link as a candidate key, and retain the title and publication date for display. Do not assume the first item is permanently the newest item without checking its date.

Timeouts, status codes, and safe retries

  • Timeout: set separate connect and read limits, as in timeout=(10, 30). Retry sparingly with backoff rather than creating a tight loop.
  • HTTP 4xx: check the URL, required parameters, and whether access is allowed. Repeating the same request will not fix a permission or malformed-URL error.
  • HTTP 5xx: record the status and retry later with bounded exponential backoff.
  • Empty item list: save the response and inspect whether it is an error page, a changed XML shape, or a feed with no current entries.
  • Encoding errors: pass response bytes, not a prematurely decoded string, so the XML declaration can guide parsing.

Simple bounded retry pattern

import time
import requests

for attempt in range(3):
    try:
        response = requests.get(FEED_URL, timeout=(10, 30))
        response.raise_for_status()
        break
    except requests.RequestException:
        if attempt == 2:
            raise
        time.sleep(2 ** attempt)

Retries should be limited to transient failures. Add a cache and a schedule appropriate to your application instead of polling continuously.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Testing and operating the extractor

Test parsing without the network

Store a representative XML fixture and test cases with a missing title, missing date, empty feed, and malformed date. This separates parser regressions from network availability.

Log useful diagnostics

  • Request timestamp and feed identifier, not sensitive credentials.
  • HTTP status and response byte count.
  • Number of item elements found.
  • Count of records missing title, link, or date.
  • Retry count and final exception type.

Do not log full responses by default if they may contain content you do not need to retain.

Dates and time zones

pubDate is commonly formatted as an RFC-style date. Parse it with a date parser that preserves its offset, as shown above, and convert to your reporting time zone only at the presentation boundary.

What Beautiful Soup does not do

  • It does not discover a feed, authenticate to Google News, or provide a Google News API.
  • It does not guarantee that an observed feed URL, field, item count, or ordering remains unchanged.
  • It does not replace responsible HTTP behavior, caching, error handling, or compliance with applicable terms.

The division of responsibility is simple: an HTTP client obtains a document, Beautiful Soup parses that document, and your application decides how to store, filter, display, and refresh the records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your broader workflow also needs screenshots of article pages, ScreenshotNeo provides a one-request website screenshot API; it is separate from RSS parsing and does not turn Google News into an API. Cookie and consent banners are accepted and removed before capture, along with more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing result in headers.

Call the API with the URL you want to capture:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

See the parameter reference and options in the ScreenshotNeo documentation. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently Asked Questions

Does Beautiful Soup provide a Google News API?

No. It parses HTML or XML that your code has already downloaded. Your HTTP client handles the request, while Google’s Feedfetcher documentation describes Google’s own user-triggered retrieval service rather than a stable public API for third-party scripts.

Why should I use the XML parser instead of an HTML parser?

RSS is XML. Calling BeautifulSoup(xml_bytes, "xml") selects XML-aware parsing and makes the intended document type explicit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I rely on a fixed number of Google News items?

No supported fixed item limit is established here. Treat item count, ordering, fields, and URL conventions as changeable.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.