You can scrape Hacker News with Requests and BeautifulSoup. Requests downloads the page, BeautifulSoup turns the markup into a searchable tree, and you pull out the story rows. But if your goal is Hacker News data rather than practice with HTML, use the official Hacker News API instead. It is public, read-only and backed by Firebase. Y Combinator launched it in 2014 so that apps built on scraping would have a stable alternative.
This article builds both: a BeautifulSoup scraper that teaches the mechanics, and an API collector you should prefer for real work.
Should you use the Hacker News API or scrape the website?
Kevin Hale, then a Y Combinator partner, explained the reason for the API in the October 7, 2014 announcement: “Because there are a lot of apps and projects out there that rely on scraping the site to access the data inside it, we decided it would be best to release a proper API and give everyone time to convert their code before we launch any new HTML.”
| Axis | Official API | HTML scraping |
|---|---|---|
| Data shape | Structured JSON records and ID lists | Markup you must parse |
| Maintenance | Documented, versioned endpoints (/v0/) |
Selectors depend on the page’s markup and on your parser |
| Request pattern | List endpoints return only IDs, so you fetch each item separately | One page fetch yields many story rows |
| Best for | Collecting HN data | Learning to parse HTML, or targets with no API |
The API’s one cost is the number of requests. The scraper’s cost is fragility: when the markup changes, your selectors break without warning.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minute#1 Best Overall
Setup
- Create a virtual environment:
python -m venv venv, then activate it. - Install the libraries:
pip install requests beautifulsoup4.
How to scrape Hacker News with BeautifulSoup
The workflow has six stages: request the page with a timeout, check the status, parse with an explicit parser, inspect the markup, handle missing values, and emit structured results.
1. Fetch the page safely
Requests applies no timeout unless you set one, so a stalled server can hang your script forever. Pass timeout, then call raise_for_status() so 4xx and 5xx responses raise an exception instead of being parsed as if they were content.
Rank #2
import requests
URL = "https://news.ycombinator.com/"
HEADERS = {"User-Agent": "learning-scraper/0.1 (contact: [email protected])"}
response = requests.get(URL, headers=HEADERS, timeout=10)
response.raise_for_status()
html = response.text
2. Parse with a named parser
BeautifulSoup builds a tree from HTML or XML. Different parser libraries can build different trees from malformed markup, so state which one you use. Python’s built-in html.parser needs no extra install.
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
3. Inspect the markup before choosing selectors
Open the page in your browser, right-click a story title and choose Inspect. Note which element wraps each story, which element holds the title link, and where the score and author appear (often in a separate row below the title). Selectors are tied to the markup you see today, so keep them in one place where they are easy to change. The values below are starting points to confirm against the live page, not guaranteed to work.
Recommended Free Tools
4. Extract rows and tolerate missing fields
find_all() searches a tag’s descendants for matching tags and attributes. Stories such as job posts may lack a score or author, so never assume an element exists.
# Selectors: verify these in your browser's inspector.
ROW_CLASS = "athing"
TITLE_SELECTOR = "span.titleline a"
def parse_stories(soup):
stories = []
for row in soup.find_all("tr", class_=ROW_CLASS):
link = row.select_one(TITLE_SELECTOR)
if link is None:
continue # layout differs or row is not a story
meta = row.find_next_sibling("tr")
score_tag = meta.select_one("span.score") if meta else None
author_tag = meta.select_one("a.hnuser") if meta else None
stories.append({
"id": row.get("id"),
"title": link.get_text(strip=True),
"url": link.get("href"),
"score": score_tag.get_text(strip=True) if score_tag else None,
"author": author_tag.get_text(strip=True) if author_tag else None,
})
return stories
Some links are relative (for example, Ask HN posts point to an internal item page), so resolve them with urllib.parse.urljoin(URL, href) if you need absolute addresses.
5. Output the results
import json
stories = parse_stories(soup)
print(json.dumps(stories[:5], indent=2))
If the list comes back empty, your selectors no longer match. Check that the response is the page you expected, then re-inspect the markup and update the constants.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.The better route: collecting stories from the API
Endpoints such as /v0/topstories and /v0/newstories return arrays of item IDs, not full stories. Per the API documentation, top and new lists hold up to 500 IDs, and the latest Ask, Show and job lists up to 200. You then fetch each record from /v0/item/<id>.json. Item fields include title, url, score, by (author), time (Unix timestamp), kids (comment IDs) and, on stories and polls, descendants (comment count).
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Best Value
import requests
BASE = "https://hacker-news.firebaseio.com/v0"
def get_json(session, path):
r = session.get(f"{BASE}/{path}", timeout=10)
r.raise_for_status()
return r.json()
def top_stories(limit=30):
with requests.Session() as s:
ids = get_json(s, "topstories.json")[:limit]
stories = []
for item_id in ids:
try:
item = get_json(s, f"item/{item_id}.json")
except requests.RequestException:
continue # skip failures; retry if you need completeness
if not item or item.get("deleted") or item.get("dead"):
continue
stories.append({
"id": item["id"],
"title": item.get("title"),
"url": item.get("url"), # absent on text posts
"score": item.get("score"),
"author": item.get("by"),
"time": item.get("time"),
"comments": item.get("descendants", 0),
})
return stories
Notes on this design:
- Ignore unknown fields. The documentation says clients should gracefully handle additional fields they don’t expect and simply ignore them. Using
.get()and picking only what you need does this. - Missing values are normal. Text posts such as Ask HN have no
url, and an item lookup can returnnull. - Be considerate. The documentation described no rate limit when it was written. That is not a guarantee about future or heavy use, so cap your batch sizes and add a short delay if you loop often.
- Convert timestamps with
datetime.fromtimestamp(item["time"], tz=timezone.utc).
When HTML parsing is still the right tool
BeautifulSoup is the right choice when you want to practice parsing and searching HTML, or when a site you care about offers no suitable API. Everything above carries over: explicit timeout, status check, named parser, selectors isolated in one place, and defensive handling of absent elements.
Optional further reading
Al Sweigart’s Automate the Boring Stuff with Python (3rd edition, No Starch Press) has a chapter titled “Web Scraping.” It is a general scraping resource, not specific to Hacker News, and print availability may change.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




