October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
AI agents

Hacker News Scraping APIs for AI Agents: Use the Official API

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For an AI agent that needs public Hacker News stories, comments, or updates, start with Hacker News’s official Firebase-backed API—not HTML scraping. It returns structured records and item IDs, with story-list endpoints for discovery and an updates endpoint for changed items and profiles. Use the Algolia-powered Hacker News interface when the task is search, and verify its coverage and freshness for your use case.

Is there an official Hacker News API?

Yes. Hacker News documents a public, Firebase-backed API at https://hacker-news.firebaseio.com/v0/. Its documentation describes public data as available in near real time. This is an API, not a requirement to scrape rendered HTML pages, and it is the most direct starting point for agents that need HN records and discussions.

The API documentation currently says there is no rate limit. Treat that as the documentation’s present statement, not a permanent service guarantee; recheck the official page before deploying a high-volume client. The documentation also warns that v0 may change and asks clients to tolerate fields they do not recognize. Parse the fields you need, and avoid rejecting an otherwise valid item because an extra field appears.

How the data is organized

HN records are items identified by integer IDs. Documented item types are job, story, comment, poll, and pollopt. Depending on the record, fields can include an author, creation time in Unix time, HTML text, parent ID, child IDs in kids, URL, score, title, poll parts, and descendant count. See the official API documentation for the current schema.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Discovery endpoints usually return arrays of IDs rather than complete stories. Fetch each item record separately. A comment may identify its parent, while a story or comment may contain kids listing child IDs. To collect a discussion, follow these links recursively as needed; a single story record does not contain the entire comment tree. HN notes that determining comment totals may require traversing that tree.

Which endpoint should an agent use?

Need Endpoint What it provides
Highest-ranked stories /v0/topstories Story IDs; top and new lists can hold up to 500 IDs, according to HN’s documentation.
Recently submitted stories /v0/newstories Story IDs; top and new lists can hold up to 500 IDs, according to HN’s documentation.
Best stories /v0/beststories IDs for the best-story list.
Ask HN /v0/askstories IDs for up to 200 of the latest Ask stories.
Show HN /v0/showstories IDs for up to 200 of the latest Show stories.
Jobs /v0/jobstories IDs for up to 200 of the latest job stories.
Changed items and profiles /v0/updates Lists of changed item IDs and profile names.
Current maximum item ID /v0/maxitem The current maximum item ID.
One record /v0/item/<id>.json A single item record, such as a story or comment.
One public profile /v0/user/<username>.json A profile for a user with public activity.

HN’s documentation says user records are available only for people with public activity, such as submitting stories or commenting. A profile may include creation time, karma, an optional HTML self-description, and submitted item IDs. Do not assume every username has a retrievable public profile.

Fetch stories and comments with a small client

This Python example discovers the latest story IDs, fetches each story, and optionally walks its comment tree. It uses only the Python standard library. The traversal is bounded by a maximum number of comments so an agent does not accidentally retrieve an unbounded discussion; increase or remove that limit only when the workflow requires it.

import json
import time
from urllib.error import HTTPError, URLError
from urllib.request import urlopen

API = "https://hacker-news.firebaseio.com/v0"

def get_json(path):
    with urlopen(f"{API}/{path}.json", timeout=20) as response:
        return json.load(response)

def get_item(item_id):
    return get_json(f"item/{item_id}")

def collect_comments(root_ids, max_comments=100):
    comments = []
    pending = list(root_ids)
    seen = set()
    while pending and len(comments) < max_comments:
        item_id = pending.pop()
        if item_id in seen:
            continue
        seen.add(item_id)
        item = get_item(item_id)
        if not item:
            continue
        comments.append(item)
        pending.extend(item.get("kids", []))
    return comments

try:
    story_ids = get_json("newstories")[:10]
    for story_id in story_ids:
        story = get_item(story_id)
        if not story:
            continue
        print(json.dumps({
            "id": story.get("id"),
            "title": story.get("title"),
            "url": story.get("url"),
            "by": story.get("by"),
            "time": story.get("time"),
            "score": story.get("score"),
        }, ensure_ascii=False))
        # Uncomment to fetch a bounded set of descendants:
        # discussion = collect_comments(story.get("kids", []), max_comments=100)
except (HTTPError, URLError, TimeoutError, json.JSONDecodeError) as exc:
    raise SystemExit(f"Hacker News API request failed: {exc}")

The example fetches records one at a time for clarity. A production agent should limit concurrency, set timeouts, handle transient failures, and cache records it has already retrieved. Avoid requesting the same item repeatedly during one run. The API’s current no-rate-limit statement does not eliminate the need for responsible request behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a reliable update loop

For ongoing monitoring, use /v0/updates to discover changed item IDs and profile names, then fetch the corresponding records. Do not infer that every new story can be found by incrementing the ID: IDs identify items, but a story-list endpoint or updates feed is the appropriate discovery mechanism. Store the last successfully processed state in your own system so a restart does not force a full refetch.

  1. At startup, request the list that matches the task, such as /v0/newstories or /v0/askstories.
  2. Fetch each relevant ID from /v0/item/<id>.json, and persist the record or the subset of fields your agent needs.
  3. On later polling cycles, request /v0/updates, then fetch changed records and profiles of interest.
  4. When a changed story’s discussion matters, inspect its kids and fetch descendants the agent has not processed.
  5. Make processing idempotent: the same ID may be encountered in more than one discovery or update pass.

The documented API is described as near real time, but the cited documentation does not establish a delivery-time guarantee. Choose a polling interval based on how fresh your application needs data to be, and verify behavior in your own workflow rather than promising users a fixed latency.

Use search when retrieval is not enough

The Firebase API is a good fit for fetching known records, following comment trees, and discovering updates. It is not documented here as a full-text query interface over stories and comments. For search-oriented work, Hacker News also has an Algolia-powered interface at hn.algolia.com/api.

The interface’s page depended on JavaScript during review, so specific endpoint parameters, historical retention depth, quotas, and completeness are not established here. Treat the search index as a distinct, separately maintained layer: before relying on it for an agent, check whether its indexed data is fresh and deep enough for the intended query.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Algolia’s developer overview describes search APIs, indexing, and search tooling for applications: Algolia developer overview. That makes hosted search infrastructure a possible option if your team collects HN records and needs its own retrieval behavior. It does not establish pricing or quotas for the Hacker News search interface. Review applicable service terms; Algolia’s terms page says it was last updated January 12, 2026: Algolia terms.

Choose the route by task

Question Official Firebase API Algolia-powered HN search
Best fit Structured record retrieval, story-list discovery, and update checking. Search-oriented discovery through a separate search interface.
Data shape Item IDs and records connected through parent and child IDs. Query and indexing behavior; specific parameters are not established here.
Freshness and coverage HN describes its public data as near real time. Index freshness and historical depth should be verified for the workflow.
Operational work Fetch linked records and, if needed, maintain an agent-side index. Use the hosted search layer; check the terms and limits that apply.
Current quotas or price for this task HN documentation currently states there is no rate limit; policies can change. Not established for the specific HN search interface.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common implementation problems

  • You received IDs but no titles or text. Discovery lists return IDs. Fetch each record with /v0/item/<id>.json.
  • A story’s comment count does not match comments fetched. The discussion is a linked tree, not a flat field in the story record. Follow kids recursively; HN notes that comment totals may require traversing the tree.
  • A requested user profile is missing. HN documents profiles for users with public activity. Do not treat every username as guaranteed to have a public record.
  • Your parser breaks after a schema change. HN says v0 may change and asks clients to tolerate unknown fields. Ignore fields you do not use and validate required fields individually.
  • Search misses a recent item or old discussion. The HN search index’s freshness and historical depth are not established here. Verify coverage for the relevant time period and use official item retrieval when you already have IDs.
  • Requests stall or fail intermittently. Set a finite timeout, retry transient network errors with backoff, and record the item ID that failed. Avoid tight retry loops; resume from saved work rather than restarting a large retrieval.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not an HN data API; it is an alternative when an AI agent needs a rendered website capture rather than structured Hacker News records. A single GET request returns a screenshot or PDF, and its options include waiting for a selector, delay, or network idle, using custom headers or cookies, and capturing full pages.

Example request, with details in the ScreenshotNeo documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://hn.algolia.com -o shot.webp
  • Cookie banners are accepted and removed, and known consent platforms, newsletter popups, and chat widgets can be removed before the capture; each step can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents using Claude, Cursor, or another MCP client.
  • The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can an AI agent scrape Hacker News without parsing HTML?

Yes. The official Firebase-backed API returns structured item records and discovery IDs, so HTML parsing is unnecessary for those public-data tasks.

Does the Hacker News API include every comment in a story response?

No. Comment relationships are represented with parent and child IDs; fetch linked descendants when the agent needs the discussion.

Is Algolia the official Hacker News data API?

No. It is a separate search-oriented interface, distinct from the official Firebase API for records and updates.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.