October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

Building a Hacker News Scraper with Python and BeautifulSoup

Learn how to parse Hacker News HTML with Requests and BeautifulSoup, and when the official Firebase API is the sturdier choice.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can scrape Hacker News with Requests and BeautifulSoup. Requests downloads the page, BeautifulSoup turns the markup into a searchable tree, and you pull out the story rows. But if your goal is Hacker News data rather than practice with HTML, use the official Hacker News API instead. It is public, read-only and backed by Firebase. Y Combinator launched it in 2014 so that apps built on scraping would have a stable alternative.

This article builds both: a BeautifulSoup scraper that teaches the mechanics, and an API collector you should prefer for real work.

Should you use the Hacker News API or scrape the website?

Kevin Hale, then a Y Combinator partner, explained the reason for the API in the October 7, 2014 announcement: “Because there are a lot of apps and projects out there that rely on scraping the site to access the data inside it, we decided it would be best to release a proper API and give everyone time to convert their code before we launch any new HTML.”

Axis Official API HTML scraping
Data shape Structured JSON records and ID lists Markup you must parse
Maintenance Documented, versioned endpoints (/v0/) Selectors depend on the page’s markup and on your parser
Request pattern List endpoints return only IDs, so you fetch each item separately One page fetch yields many story rows
Best for Collecting HN data Learning to parse HTML, or targets with no API

The API’s one cost is the number of requests. The scraper’s cost is fragility: when the markup changes, your selectors break without warning.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Setup

  1. Create a virtual environment: python -m venv venv, then activate it.
  2. Install the libraries: pip install requests beautifulsoup4.

How to scrape Hacker News with BeautifulSoup

The workflow has six stages: request the page with a timeout, check the status, parse with an explicit parser, inspect the markup, handle missing values, and emit structured results.

1. Fetch the page safely

Requests applies no timeout unless you set one, so a stalled server can hang your script forever. Pass timeout, then call raise_for_status() so 4xx and 5xx responses raise an exception instead of being parsed as if they were content.

import requests

URL = "https://news.ycombinator.com/"
HEADERS = {"User-Agent": "learning-scraper/0.1 (contact: [email protected])"}

response = requests.get(URL, headers=HEADERS, timeout=10)
response.raise_for_status()
html = response.text

2. Parse with a named parser

BeautifulSoup builds a tree from HTML or XML. Different parser libraries can build different trees from malformed markup, so state which one you use. Python’s built-in html.parser needs no extra install.

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")

3. Inspect the markup before choosing selectors

Open the page in your browser, right-click a story title and choose Inspect. Note which element wraps each story, which element holds the title link, and where the score and author appear (often in a separate row below the title). Selectors are tied to the markup you see today, so keep them in one place where they are easy to change. The values below are starting points to confirm against the live page, not guaranteed to work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Extract rows and tolerate missing fields

find_all() searches a tag’s descendants for matching tags and attributes. Stories such as job posts may lack a score or author, so never assume an element exists.

# Selectors: verify these in your browser's inspector.
ROW_CLASS = "athing"
TITLE_SELECTOR = "span.titleline a"

def parse_stories(soup):
    stories = []
    for row in soup.find_all("tr", class_=ROW_CLASS):
        link = row.select_one(TITLE_SELECTOR)
        if link is None:
            continue  # layout differs or row is not a story

        meta = row.find_next_sibling("tr")
        score_tag = meta.select_one("span.score") if meta else None
        author_tag = meta.select_one("a.hnuser") if meta else None

        stories.append({
            "id": row.get("id"),
            "title": link.get_text(strip=True),
            "url": link.get("href"),
            "score": score_tag.get_text(strip=True) if score_tag else None,
            "author": author_tag.get_text(strip=True) if author_tag else None,
        })
    return stories

Some links are relative (for example, Ask HN posts point to an internal item page), so resolve them with urllib.parse.urljoin(URL, href) if you need absolute addresses.

5. Output the results

import json

stories = parse_stories(soup)
print(json.dumps(stories[:5], indent=2))

If the list comes back empty, your selectors no longer match. Check that the response is the page you expected, then re-inspect the markup and update the constants.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

The better route: collecting stories from the API

Endpoints such as /v0/topstories and /v0/newstories return arrays of item IDs, not full stories. Per the API documentation, top and new lists hold up to 500 IDs, and the latest Ask, Show and job lists up to 200. You then fetch each record from /v0/item/<id>.json. Item fields include title, url, score, by (author), time (Unix timestamp), kids (comment IDs) and, on stories and polls, descendants (comment count).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

BASE = "https://hacker-news.firebaseio.com/v0"

def get_json(session, path):
    r = session.get(f"{BASE}/{path}", timeout=10)
    r.raise_for_status()
    return r.json()

def top_stories(limit=30):
    with requests.Session() as s:
        ids = get_json(s, "topstories.json")[:limit]
        stories = []
        for item_id in ids:
            try:
                item = get_json(s, f"item/{item_id}.json")
            except requests.RequestException:
                continue  # skip failures; retry if you need completeness
            if not item or item.get("deleted") or item.get("dead"):
                continue
            stories.append({
                "id": item["id"],
                "title": item.get("title"),
                "url": item.get("url"),  # absent on text posts
                "score": item.get("score"),
                "author": item.get("by"),
                "time": item.get("time"),
                "comments": item.get("descendants", 0),
            })
        return stories

Notes on this design:

  • Ignore unknown fields. The documentation says clients should gracefully handle additional fields they don’t expect and simply ignore them. Using .get() and picking only what you need does this.
  • Missing values are normal. Text posts such as Ask HN have no url, and an item lookup can return null.
  • Be considerate. The documentation described no rate limit when it was written. That is not a guarantee about future or heavy use, so cap your batch sizes and add a short delay if you loop often.
  • Convert timestamps with datetime.fromtimestamp(item["time"], tz=timezone.utc).

When HTML parsing is still the right tool

BeautifulSoup is the right choice when you want to practice parsing and searching HTML, or when a site you care about offers no suitable API. Everything above carries over: explicit timeout, status check, named parser, selectors isolated in one place, and defensive handling of absent elements.

Optional further reading

Al Sweigart’s Automate the Boring Stuff with Python (3rd edition, No Starch Press) has a chapter titled “Web Scraping.” It is a general scraping resource, not specific to Hacker News, and print availability may change.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.